# All 150 crawlers

> Every crawler in the index — 150 entries across 74 operators, each checked against the operator's own documentation. One page, one JSON record and one markdown file per crawler.

```
curl -s https://www.pathwren.workers.dev/data/agents.json | jq '.crawlers[].robots_token'
```

| Crawler | Operator | Category | robots.txt token |
|---|---|---|---|
| [AdsBot-Google](/crawler/adsbot-google.html) | Google | Tools and frameworks | `AdsBot-Google` |
| [AdsBot-Google-Mobile](/crawler/adsbot-google-mobile.html) | Google | Tools and frameworks | `AdsBot-Google-Mobile` |
| [AdsBot-Google-Mobile-Apps](/crawler/adsbot-google-mobile-apps.html) | Google | Tools and frameworks | `AdsBot-Google-Mobile-Apps` |
| [AhrefsBot](/crawler/ahrefsbot.html) | Ahrefs | SEO and backlink crawlers | `AhrefsBot` |
| [AhrefsSiteAudit](/crawler/ahrefssiteaudit.html) | Ahrefs | SEO and backlink crawlers | `AhrefsSiteAudit` |
| [AI2Bot](/crawler/ai2bot.html) | Allen Institute for AI | Corpus and dataset builders | `AI2Bot` |
| [Ai2Bot-Dolma](/crawler/ai2bot-dolma.html) | Allen Institute for AI | Corpus and dataset builders | `Ai2Bot-Dolma` |
| [aiHitBot](/crawler/aihitbot.html) | aiHit | Corpus and dataset builders | `aiHitBot` |
| [AIWebIndex](/crawler/aiwebindex.html) | Lyrenth | AI search crawlers | `AIWebIndex` |
| [Amazonbot](/crawler/amazonbot.html) | Amazon | AI search crawlers | `Amazonbot` |
| [Andibot](/crawler/andibot.html) | Andi | AI search crawlers | `Andibot` |
| [Anomura](/crawler/anomura.html) | Direqt | AI search crawlers | `Anomura` |
| [anthropic-ai](/crawler/anthropic-ai.html) | Anthropic | AI training crawlers | `anthropic-ai` |
| [APIs-Google](/crawler/apis-google.html) | Google | Tools and frameworks | `APIs-Google` |
| [Applebot](/crawler/applebot.html) | Apple | Search engines | `Applebot` |
| [Applebot-Extended](/crawler/applebot-extended.html) | Apple | AI training crawlers | `Applebot-Extended` |
| [archive.org_bot](/crawler/archive-org-bot.html) | Internet Archive | Archivers | `archive.org_bot` |
| [atlassian-bot](/crawler/atlassian-bot.html) | Atlassian | AI search crawlers | `atlassian-bot` |
| [AwarioRssBot](/crawler/awariorssbot.html) | Awario | Corpus and dataset builders | `AwarioRssBot` |
| [AwarioSmartBot](/crawler/awariosmartbot.html) | Awario | Corpus and dataset builders | `AwarioSmartBot` |
| [Baiduspider](/crawler/baiduspider.html) | Baidu | Search engines | `Baiduspider` |
| [Barkrowler](/crawler/barkrowler.html) | Babbar | SEO and backlink crawlers | `barkrowler` |
| [bedrockbot](/crawler/bedrockbot.html) | Amazon | AI search crawlers | `bedrockbot` |
| [bingbot](/crawler/bingbot.html) | Microsoft | Search engines | `bingbot` |
| [Bytespider](/crawler/bytespider.html) | ByteDance | AI training crawlers | `Bytespider` |
| [CCBot](/crawler/ccbot.html) | Common Crawl | Corpus and dataset builders | `CCBot` |
| [ChatGPT Agent](/crawler/chatgpt-agent.html) | OpenAI | User-triggered fetchers | `ChatGPT-User` |
| [ChatGPT-User](/crawler/chatgpt-user.html) | OpenAI | User-triggered fetchers | `ChatGPT-User` |
| [Claude-SearchBot](/crawler/claude-searchbot.html) | Anthropic | AI search crawlers | `Claude-SearchBot` |
| [Claude-User](/crawler/claude-user.html) | Anthropic | User-triggered fetchers | `Claude-User` |
| [Claude-Web](/crawler/claude-web.html) | Anthropic | AI search crawlers | `Claude-Web` |
| [ClaudeBot](/crawler/claudebot.html) | Anthropic | AI training crawlers | `ClaudeBot` |
| [Cloudflare-AutoRAG](/crawler/cloudflare-autorag.html) | Cloudflare | AI search crawlers | `Cloudflare-AutoRAG` |
| [cohere-ai](/crawler/cohere-ai.html) | Cohere | User-triggered fetchers | `cohere-ai` |
| [cohere-training-data-crawler](/crawler/cohere-training-data-crawler.html) | Cohere | AI training crawlers | `cohere-training-data-crawler` |
| [Cotoyogi](/crawler/cotoyogi.html) | ROIS-DS | AI training crawlers | `Cotoyogi` |
| [Crawl4AI](/crawler/crawl4ai.html) | Crawl4AI project | Tools and frameworks | `Crawl4AI` |
| [Crawlspace](/crawler/crawlspace.html) | Crawlspace | Tools and frameworks | `Crawlspace` |
| [DataForSeoBot](/crawler/dataforseobot.html) | DataForSEO | SEO and backlink crawlers | `DataForSeoBot` |
| [Diffbot](/crawler/diffbot.html) | Diffbot | Corpus and dataset builders | `Diffbot` |
| [DotBot](/crawler/dotbot.html) | Moz | SEO and backlink crawlers | `dotbot` |
| [DuckAssistBot](/crawler/duckassistbot.html) | DuckDuckGo | AI search crawlers | `DuckAssistBot` |
| [DuckDuckBot](/crawler/duckduckbot.html) | DuckDuckGo | Search engines | `DuckDuckBot` |
| [EchoboxBot](/crawler/echoboxbot.html) | Echobox | Corpus and dataset builders | `EchoboxBot` |
| [ExaSearchBot](/crawler/exasearchbot.html) | Exa | AI search crawlers | `ExaSearchBot` |
| [FacebookBot](/crawler/facebookbot.html) | Meta | AI training crawlers | `FacebookBot` |
| [facebookexternalhit](/crawler/facebookexternalhit.html) | Meta | Link preview fetchers | `facebookexternalhit` |
| [Factset_spyderbot](/crawler/factset-spyderbot.html) | FactSet | AI training crawlers | `Factset_spyderbot` |
| [FeedFetcher-Google](/crawler/feedfetcher-google.html) | Google | Tools and frameworks | `FeedFetcher-Google` |
| [FirecrawlAgent](/crawler/firecrawlagent.html) | Firecrawl | Tools and frameworks | `FirecrawlAgent` |
| [Google-Agent](/crawler/google-agent.html) | Google | User-triggered fetchers | `Google-Agent` |
| [Google-CloudVertexBot](/crawler/google-cloudvertexbot.html) | Google | AI search crawlers | `Google-CloudVertexBot` |
| [Google-CWS](/crawler/google-cws.html) | Google | Tools and frameworks | `Google-CWS` |
| [Google-Extended](/crawler/google-extended.html) | Google | AI training crawlers | `Google-Extended` |
| [Google-GeminiNotebook](/crawler/google-gemininotebook.html) | Google | User-triggered fetchers | `Google-GeminiNotebook` |
| [Google-InspectionTool](/crawler/google-inspectiontool.html) | Google | Tools and frameworks | `Google-InspectionTool` |
| [Google-Pinpoint](/crawler/google-pinpoint.html) | Google | User-triggered fetchers | `Google-Pinpoint` |
| [Google-Read-Aloud](/crawler/google-read-aloud.html) | Google | User-triggered fetchers | `Google-Read-Aloud` |
| [Google-Safety](/crawler/google-safety.html) | Google | Tools and frameworks | `Google-Safety` |
| [Google-Site-Verification](/crawler/google-site-verification.html) | Google | Tools and frameworks | `Google-Site-Verification` |
| [Googlebot](/crawler/googlebot.html) | Google | Search engines | `Googlebot` |
| [Googlebot-Image](/crawler/googlebot-image.html) | Google | Search engines | `Googlebot-Image` |
| [Googlebot-News](/crawler/googlebot-news.html) | Google | Search engines | `Googlebot-News` |
| [Googlebot-Video](/crawler/googlebot-video.html) | Google | Search engines | `Googlebot-Video` |
| [GoogleMessages](/crawler/googlemessages.html) | Google | Link preview fetchers | `GoogleMessages` |
| [GoogleOther](/crawler/googleother.html) | Google | AI training crawlers | `GoogleOther` |
| [GoogleOther-Image](/crawler/googleother-image.html) | Google | AI training crawlers | `GoogleOther-Image` |
| [GoogleOther-Video](/crawler/googleother-video.html) | Google | AI training crawlers | `GoogleOther-Video` |
| [GoogleProducer](/crawler/googleproducer.html) | Google | Tools and frameworks | `GoogleProducer` |
| [GPTBot](/crawler/gptbot.html) | OpenAI | AI training crawlers | `GPTBot` |
| [ia_archiver](/crawler/ia-archiver.html) | Internet Archive | Archivers | `ia_archiver` |
| [ICC-Crawler](/crawler/icc-crawler.html) | NICT | AI training crawlers | `ICC-Crawler` |
| [ImagesiftBot](/crawler/imagesiftbot.html) | Hive AI | Corpus and dataset builders | `ImagesiftBot` |
| [img2dataset](/crawler/img2dataset.html) | LAION / img2dataset | Corpus and dataset builders | `img2dataset` |
| [ISSCyberRiskCrawler](/crawler/isscyberriskcrawler.html) | ISS Corporate Solutions | AI training crawlers | `ISSCyberRiskCrawler` |
| [Kagibot](/crawler/kagibot.html) | Kagi | Search engines | `Kagibot` |
| [KlaviyoAIBot](/crawler/klaviyoaibot.html) | Klaviyo | AI search crawlers | `KlaviyoAIBot` |
| [LAIONDownloader](/crawler/laiondownloader.html) | LAION / img2dataset | Corpus and dataset builders | `LAIONDownloader` |
| [Lightpanda](/crawler/lightpanda.html) | Lightpanda | Tools and frameworks | `Lightpanda` |
| [Linguee Bot](/crawler/linguee-bot.html) | Linguee | AI training crawlers | `Linguee Bot` |
| [Mediapartners-Google](/crawler/mediapartners-google.html) | Google | Tools and frameworks | `Mediapartners-Google` |
| [meta-externalagent](/crawler/meta-externalagent.html) | Meta | AI training crawlers | `meta-externalagent` |
| [meta-externalfetcher](/crawler/meta-externalfetcher.html) | Meta | User-triggered fetchers | `meta-externalfetcher` |
| [Meta-WebIndexer](/crawler/meta-webindexer.html) | Meta | AI search crawlers | `Meta-WebIndexer` |
| [MistralAI-User](/crawler/mistralai-user.html) | Mistral AI | User-triggered fetchers | `MistralAI-User` |
| [MJ12bot](/crawler/mj12bot.html) | Majestic | SEO and backlink crawlers | `MJ12bot` |
| [MojeekBot](/crawler/mojeekbot.html) | Mojeek | Search engines | `MojeekBot` |
| [OAI-SearchBot](/crawler/oai-searchbot.html) | OpenAI | AI search crawlers | `OAI-SearchBot` |
| [omgili](/crawler/omgili.html) | Webz.io | Corpus and dataset builders | `omgili` |
| [omgilibot](/crawler/omgilibot.html) | Webz.io | Corpus and dataset builders | `omgilibot` |
| [Panscient](/crawler/panscient.html) | Panscient | Corpus and dataset builders | `panscient.com` |
| [Perplexity-User](/crawler/perplexity-user.html) | Perplexity | User-triggered fetchers | `Perplexity-User` |
| [PerplexityBot](/crawler/perplexitybot.html) | Perplexity | AI search crawlers | `PerplexityBot` |
| [PetalBot](/crawler/petalbot.html) | Huawei | Search engines | `PetalBot` |
| [PhindBot](/crawler/phindbot.html) | Phind | AI search crawlers | `PhindBot` |
| [Pinterestbot](/crawler/pinterestbot.html) | Pinterest | Search engines | `Pinterestbot` |
| [Poseidon Research Crawler](/crawler/poseidon-research-crawler.html) | Poseidon Research | AI training crawlers | `Poseidon Research Crawler` |
| [QualifiedBot](/crawler/qualifiedbot.html) | Qualified | AI search crawlers | `QualifiedBot` |
| [QuillBot](/crawler/quillbot.html) | QuillBot | AI training crawlers | `QuillBot` |
| [Qwantbot](/crawler/qwantbot.html) | Qwant | Search engines | `Qwantbot` |
| [Qwantbot-news](/crawler/qwantbot-news.html) | Qwant | Search engines | `Qwantbot-news` |
| [Reflectionbot](/crawler/reflectionbot.html) | Reflection AI | AI training crawlers | `Reflectionbot` |
| [rogerbot](/crawler/rogerbot.html) | Moz | SEO and backlink crawlers | `rogerbot` |
| [SBIntuitionsBot](/crawler/sbintuitionsbot.html) | SB Intuitions | AI training crawlers | `SBIntuitionsBot` |
| [Scrapy](/crawler/scrapy.html) | Scrapy project | Tools and frameworks | `Scrapy` |
| [Screaming Frog SEO Spider](/crawler/screaming-frog-seo-spider.html) | Screaming Frog | Tools and frameworks | `Screaming Frog SEO Spider` |
| [SemrushBot](/crawler/semrushbot.html) | Semrush | SEO and backlink crawlers | `SemrushBot` |
| [SemrushBot-BA](/crawler/semrushbot-ba.html) | Semrush | SEO and backlink crawlers | `SemrushBot-BA` |
| [SemrushBot-ESI](/crawler/semrushbot-esi.html) | Semrush | SEO and backlink crawlers | `SemrushBot-ESI` |
| [SemrushBot-FT](/crawler/semrushbot-ft.html) | Semrush | SEO and backlink crawlers | `SemrushBot-FT` |
| [SemrushBot-OCOB](/crawler/semrushbot-ocob.html) | Semrush | AI training crawlers | `SemrushBot-OCOB` |
| [SemrushBot-SI](/crawler/semrushbot-si.html) | Semrush | SEO and backlink crawlers | `SemrushBot-SI` |
| [SemrushBot-SWA](/crawler/semrushbot-swa.html) | Semrush | SEO and backlink crawlers | `SemrushBot-SWA` |
| [SEOkicks](/crawler/seokicks.html) | SEOkicks | SEO and backlink crawlers | `SEOkicks` |
| [serpstatbot](/crawler/serpstatbot.html) | Serpstat | SEO and backlink crawlers | `serpstatbot` |
| [SeznamBot](/crawler/seznambot.html) | Seznam | Search engines | `SeznamBot` |
| [ShapBot](/crawler/shapbot.html) | Parallel | AI search crawlers | `ShapBot` |
| [Sidetrade indexer bot](/crawler/sidetrade-indexer-bot.html) | Sidetrade | AI training crawlers | `Sidetrade indexer bot` |
| [SiteAuditBot](/crawler/siteauditbot.html) | Semrush | SEO and backlink crawlers | `SiteAuditBot` |
| [Slackbot](/crawler/slackbot.html) | Slack | Link preview fetchers | `Slackbot` |
| [Slackbot-LinkExpanding](/crawler/slackbot-linkexpanding.html) | Slack | Link preview fetchers | `Slackbot-LinkExpanding` |
| [SplitSignalBot](/crawler/splitsignalbot.html) | Semrush | SEO and backlink crawlers | `SplitSignalBot` |
| [Storebot-Google](/crawler/storebot-google.html) | Google | Search engines | `Storebot-Google` |
| [TerraCotta](/crawler/terracotta.html) | Ceramic AI | AI search crawlers | `TerraCotta` |
| [Thinkbot](/crawler/thinkbot.html) | Thinkbot | Corpus and dataset builders | `Thinkbot` |
| [TikTokSpider](/crawler/tiktokspider.html) | ByteDance | AI training crawlers | `TikTokSpider` |
| [Timpibot](/crawler/timpibot.html) | Timpi | Search engines | `Timpibot` |
| [VelenPublicWebCrawler](/crawler/velenpublicwebcrawler.html) | Hunter (Velen) | Corpus and dataset builders | `VelenPublicWebCrawler` |
| [Webzio-Extended](/crawler/webzio-extended.html) | Webz.io | AI training crawlers | `Webzio-Extended` |
| [wpbot](/crawler/wpbot.html) | QuantumCloud | Tools and frameworks | `wpbot` |
| [YaK](/crawler/yak.html) | Meltwater | Corpus and dataset builders | `YaK` |
| [YandexAdditional](/crawler/yandexadditional.html) | Yandex | AI training crawlers | `YandexAdditional` |
| [YandexAdditionalBot](/crawler/yandexadditionalbot.html) | Yandex | AI training crawlers | `YandexAdditionalBot` |
| [YandexBlogs](/crawler/yandexblogs.html) | Yandex | Search engines | `YandexBlogs` |
| [YandexBot](/crawler/yandexbot.html) | Yandex | Search engines | `YandexBot` |
| [YandexCalendar](/crawler/yandexcalendar.html) | Yandex | User-triggered fetchers | `YandexCalendar` |
| [YandexComBot](/crawler/yandexcombot.html) | Yandex | Search engines | `YandexComBot` |
| [YandexDirect](/crawler/yandexdirect.html) | Yandex | Tools and frameworks | `YandexDirect` |
| [YandexFavicons](/crawler/yandexfavicons.html) | Yandex | Search engines | `YandexFavicons` |
| [YandexImages](/crawler/yandeximages.html) | Yandex | Search engines | `YandexImages` |
| [YandexMarket](/crawler/yandexmarket.html) | Yandex | Search engines | `YandexMarket` |
| [YandexMedia](/crawler/yandexmedia.html) | Yandex | Search engines | `YandexMedia` |
| [YandexMetrika](/crawler/yandexmetrika.html) | Yandex | Tools and frameworks | `YandexMetrika` |
| [YandexMobileBot](/crawler/yandexmobilebot.html) | Yandex | Search engines | `YandexMobileBot` |
| [YandexRenderResourcesBot](/crawler/yandexrenderresourcesbot.html) | Yandex | Search engines | `YandexRenderResourcesBot` |
| [YandexScreenshotBot](/crawler/yandexscreenshotbot.html) | Yandex | Tools and frameworks | `YandexScreenshotBot` |
| [YandexVideo](/crawler/yandexvideo.html) | Yandex | Search engines | `YandexVideo` |
| [YandexWebmaster](/crawler/yandexwebmaster.html) | Yandex | Tools and frameworks | `YandexWebmaster` |
| [Yeti](/crawler/yeti.html) | Naver | Search engines | `Yeti` |
| [YouBot](/crawler/youbot.html) | You.com | AI search crawlers | `YouBot` |

---

Machine copies of this listing: [/crawler/index.json](/crawler/index.json) ·
[/crawler/index.html](/crawler/index.html) · [everything here](/documents.json) ·
[what changed since your cursor](/changes.json).
`https://www.pathwren.workers.dev/crawler` serves this file to a client that ranks `text/markdown` above
`text/html`, the JSON to one that asks for `application/json`, and the page to
everybody else; the canonical is /crawler/index.html.

Rebuilt 2026-09-03. An independent, non-commercial automated project: it is run by software rather than by a person, and it says so wherever it introduces itself. It is not affiliated with, endorsed by or operated by any of the crawler operators it documents, nor by any other company. The category and cost-of-blocking fields are its own assessment and are labelled as such; every other field is cited to the operator's own documentation.
