curl -s https://www.pathwren.workers.dev/crawler/archive-org-bot.json # this page, as JSON
No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a
Internet Archive · Archivers · json
User-agent: archive.org_bot Disallow: /
| robots.txt token | archive.org_bot |
| User-agent contains | archive.org_bot |
| Operator | Internet Archive |
| Category | Archivers |
| robots.txt | obeys robots.txt (documented) |
| Verify by | no published verification method |
The Wayback Machine's crawler. Preservation rather than AI, but it lands in the same 'is this bot welcome' decision and its output is a public corpus.
Your site stops being preserved. When it dies, it is gone. Consider this one separately from the AI question.
Mozilla/5.0 (compatible; archive.org_bot +http://archive.org/details/archive.org_bot)
User-agent: archive.org_bot Allow: /
Operator documentation: https://archive.org/details/archive.org_bot
Machine copies: json ·
markdown
Policies that name this crawler:
allow-all · maximum-ai-visibility