curl -s https://www.pathwren.workers.dev/crawler/index.json # this page, as JSON curl -s https://www.pathwren.workers.dev/data/agents.json # Every crawler record in one file
No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a
Grouped by what the crawl is for. Machine copy: agents.json · agents.csv · user-agents.txt
Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| anthropic-ai | anthropic-ai | Anthropic | n-a |
| Applebot-Extended | Applebot-Extended | Apple | n-a |
| Bytespider | Bytespider | ByteDance | disputed |
| ClaudeBot | ClaudeBot | Anthropic | documented |
| cohere-training-data-crawler | cohere-training-data-crawler | Cohere | documented |
| Cotoyogi | Cotoyogi | ROIS-DS | documented |
| FacebookBot | FacebookBot | Meta | documented |
| Factset_spyderbot | Factset_spyderbot | FactSet | undocumented |
| Google-Extended | Google-Extended | n-a | |
| GoogleOther | GoogleOther | documented | |
| GoogleOther-Image | GoogleOther-Image | documented | |
| GoogleOther-Video | GoogleOther-Video | documented | |
| GPTBot | GPTBot | OpenAI | documented |
| ICC-Crawler | ICC-Crawler | NICT | documented |
| ISSCyberRiskCrawler | ISSCyberRiskCrawler | ISS Corporate Solutions | disputed |
| Linguee Bot | Linguee Bot | Linguee | disputed |
| meta-externalagent | meta-externalagent | Meta | documented |
| Poseidon Research Crawler | Poseidon Research Crawler | Poseidon Research | undocumented |
| QuillBot | QuillBot | QuillBot | undocumented |
| Reflectionbot | Reflectionbot | Reflection AI | undocumented |
| SBIntuitionsBot | SBIntuitionsBot | SB Intuitions | documented |
| SemrushBot-OCOB | SemrushBot-OCOB | Semrush | documented |
| Sidetrade indexer bot | Sidetrade indexer bot | Sidetrade | undocumented |
| TikTokSpider | TikTokSpider | ByteDance | disputed |
| Webzio-Extended | Webzio-Extended | Webz.io | documented |
| YandexAdditional | YandexAdditional | Yandex | own-token-only |
| YandexAdditionalBot | YandexAdditionalBot | Yandex | own-token-only |
Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| AIWebIndex | AIWebIndex | Lyrenth | documented |
| Amazonbot | Amazonbot | Amazon | documented |
| Andibot | Andibot | Andi | undocumented |
| Anomura | Anomura | Direqt | documented |
| atlassian-bot | atlassian-bot | Atlassian | documented |
| bedrockbot | bedrockbot | Amazon | documented |
| Claude-SearchBot | Claude-SearchBot | Anthropic | documented |
| Claude-Web | Claude-Web | Anthropic | n-a |
| Cloudflare-AutoRAG | Cloudflare-AutoRAG | Cloudflare | documented |
| DuckAssistBot | DuckAssistBot | DuckDuckGo | documented |
| ExaSearchBot | ExaSearchBot | Exa | undocumented |
| Google-CloudVertexBot | Google-CloudVertexBot | documented | |
| KlaviyoAIBot | KlaviyoAIBot | Klaviyo | documented |
| Meta-WebIndexer | Meta-WebIndexer | Meta | undocumented |
| OAI-SearchBot | OAI-SearchBot | OpenAI | documented |
| PerplexityBot | PerplexityBot | Perplexity | documented |
| PhindBot | PhindBot | Phind | undocumented |
| QualifiedBot | QualifiedBot | Qualified | undocumented |
| ShapBot | ShapBot | Parallel | documented |
| TerraCotta | TerraCotta | Ceramic AI | documented |
| YouBot | YouBot | You.com | documented |
Fetch one page because a person asked for it, right then. One human intent, one request. Blocking them produces a visible error for a real reader.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| ChatGPT Agent | ChatGPT-User | OpenAI | documented |
| ChatGPT-User | ChatGPT-User | OpenAI | documented |
| Claude-User | Claude-User | Anthropic | documented |
| cohere-ai | cohere-ai | Cohere | documented |
| Google-Agent | Google-Agent | by-design-no | |
| Google-GeminiNotebook | Google-GeminiNotebook | by-design-no | |
| Google-Pinpoint | Google-Pinpoint | by-design-no | |
| Google-Read-Aloud | Google-Read-Aloud | by-design-no | |
| meta-externalfetcher | meta-externalfetcher | Meta | documented |
| MistralAI-User | MistralAI-User | Mistral AI | documented |
| Perplexity-User | Perplexity-User | Perplexity | by-design-no |
| YandexCalendar | YandexCalendar | Yandex | own-token-only |
Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| AI2Bot | AI2Bot | Allen Institute for AI | documented |
| Ai2Bot-Dolma | Ai2Bot-Dolma | Allen Institute for AI | documented |
| aiHitBot | aiHitBot | aiHit | documented |
| AwarioRssBot | AwarioRssBot | Awario | documented |
| AwarioSmartBot | AwarioSmartBot | Awario | documented |
| CCBot | CCBot | Common Crawl | documented |
| Diffbot | Diffbot | Diffbot | documented |
| EchoboxBot | EchoboxBot | Echobox | undocumented |
| ImagesiftBot | ImagesiftBot | Hive AI | documented |
| img2dataset | img2dataset | LAION / img2dataset | documented |
| LAIONDownloader | LAIONDownloader | LAION / img2dataset | by-design-no |
| omgili | omgili | Webz.io | documented |
| omgilibot | omgilibot | Webz.io | documented |
| Panscient | panscient.com | Panscient | documented |
| Thinkbot | Thinkbot | Thinkbot | disputed |
| VelenPublicWebCrawler | VelenPublicWebCrawler | Hunter (Velen) | documented |
| YaK | YaK | Meltwater | undocumented |
Classic index-and-rank crawlers. Several also feed their operator's generative answers, which is why the AI opt-out for Google and Apple is a token rather than a block.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| Applebot | Applebot | Apple | documented |
| Baiduspider | Baiduspider | Baidu | documented |
| bingbot | bingbot | Microsoft | documented |
| DuckDuckBot | DuckDuckBot | DuckDuckGo | documented |
| Googlebot | Googlebot | documented | |
| Googlebot-Image | Googlebot-Image | documented | |
| Googlebot-News | Googlebot-News | documented | |
| Googlebot-Video | Googlebot-Video | documented | |
| Kagibot | Kagibot | Kagi | documented |
| MojeekBot | MojeekBot | Mojeek | documented |
| PetalBot | PetalBot | Huawei | documented |
| Pinterestbot | Pinterestbot | documented | |
| Qwantbot | Qwantbot | Qwant | documented |
| Qwantbot-news | Qwantbot-news | Qwant | documented |
| SeznamBot | SeznamBot | Seznam | documented |
| Storebot-Google | Storebot-Google | documented | |
| Timpibot | Timpibot | Timpi | documented |
| YandexBlogs | YandexBlogs | Yandex | documented |
| YandexBot | YandexBot | Yandex | documented |
| YandexComBot | YandexComBot | Yandex | own-token-only |
| YandexFavicons | YandexFavicons | Yandex | own-token-only |
| YandexImages | YandexImages | Yandex | documented |
| YandexMarket | YandexMarket | Yandex | documented |
| YandexMedia | YandexMedia | Yandex | documented |
| YandexMobileBot | YandexMobileBot | Yandex | own-token-only |
| YandexRenderResourcesBot | YandexRenderResourcesBot | Yandex | own-token-only |
| YandexVideo | YandexVideo | Yandex | documented |
| Yeti | Yeti | Naver | documented |
Commercial link-graph tooling. No user-facing effect either way, and usually a large share of your bot bandwidth.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| AhrefsBot | AhrefsBot | Ahrefs | documented |
| AhrefsSiteAudit | AhrefsSiteAudit | Ahrefs | documented |
| Barkrowler | barkrowler | Babbar | documented |
| DataForSeoBot | DataForSeoBot | DataForSEO | documented |
| DotBot | dotbot | Moz | documented |
| MJ12bot | MJ12bot | Majestic | documented |
| rogerbot | rogerbot | Moz | documented |
| SemrushBot | SemrushBot | Semrush | documented |
| SemrushBot-BA | SemrushBot-BA | Semrush | documented |
| SemrushBot-ESI | SemrushBot-ESI | Semrush | documented |
| SemrushBot-FT | SemrushBot-FT | Semrush | documented |
| SemrushBot-SI | SemrushBot-SI | Semrush | documented |
| SemrushBot-SWA | SemrushBot-SWA | Semrush | documented |
| SEOkicks | SEOkicks | SEOkicks | documented |
| serpstatbot | serpstatbot | Serpstat | documented |
| SiteAuditBot | SiteAuditBot | Semrush | documented |
| SplitSignalBot | SplitSignalBot | Semrush | documented |
Preservation crawlers. Their output is public and permanent, which makes them a separate decision from the AI one.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| archive.org_bot | archive.org_bot | Internet Archive | documented |
| ia_archiver | ia_archiver | Internet Archive | documented |
Read your Open Graph tags when someone shares a link. Blocking these is almost always an accident.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| facebookexternalhit | facebookexternalhit | Meta | documented |
| GoogleMessages | GoogleMessages | by-design-no | |
| Slackbot | Slackbot | Slack | documented |
| Slackbot-LinkExpanding | Slackbot-LinkExpanding | Slack | documented |
Not operators: crawling software anyone can run. The party behind the request is unknown, so treat them as a rate-limit question rather than a consent question.
| Crawler | robots.txt token | Operator | robots.txt |
|---|---|---|---|
| AdsBot-Google | AdsBot-Google | own-token-only | |
| AdsBot-Google-Mobile | AdsBot-Google-Mobile | own-token-only | |
| AdsBot-Google-Mobile-Apps | AdsBot-Google-Mobile-Apps | own-token-only | |
| APIs-Google | APIs-Google | own-token-only | |
| Crawl4AI | Crawl4AI | Crawl4AI project | undocumented |
| Crawlspace | Crawlspace | Crawlspace | documented |
| FeedFetcher-Google | FeedFetcher-Google | by-design-no | |
| FirecrawlAgent | FirecrawlAgent | Firecrawl | documented |
| Google-CWS | Google-CWS | by-design-no | |
| Google-InspectionTool | Google-InspectionTool | documented | |
| Google-Safety | Google-Safety | by-design-no | |
| Google-Site-Verification | Google-Site-Verification | by-design-no | |
| GoogleProducer | GoogleProducer | by-design-no | |
| Lightpanda | Lightpanda | Lightpanda | undocumented |
| Mediapartners-Google | Mediapartners-Google | own-token-only | |
| Scrapy | Scrapy | Scrapy project | documented |
| Screaming Frog SEO Spider | Screaming Frog SEO Spider | Screaming Frog | documented |
| wpbot | wpbot | QuantumCloud | undocumented |
| YandexDirect | YandexDirect | Yandex | own-token-only |
| YandexMetrika | YandexMetrika | Yandex | by-design-no |
| YandexScreenshotBot | YandexScreenshotBot | Yandex | own-token-only |
| YandexWebmaster | YandexWebmaster | Yandex | documented |