curl -s https://www.pathwren.workers.dev/crawler/index.json   # this page, as JSON
curl -s https://www.pathwren.workers.dev/data/agents.json     # Every crawler record in one file

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

All crawlers (150)

Grouped by what the crawl is for. Machine copy: agents.json · agents.csv · user-agents.txt

AI training crawlers (27)

Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today.

Crawlerrobots.txt tokenOperatorrobots.txt
anthropic-aianthropic-aiAnthropicn-a
Applebot-ExtendedApplebot-ExtendedApplen-a
BytespiderBytespiderByteDancedisputed
ClaudeBotClaudeBotAnthropicdocumented
cohere-training-data-crawlercohere-training-data-crawlerCoheredocumented
CotoyogiCotoyogiROIS-DSdocumented
FacebookBotFacebookBotMetadocumented
Factset_spyderbotFactset_spyderbotFactSetundocumented
Google-ExtendedGoogle-ExtendedGooglen-a
GoogleOtherGoogleOtherGoogledocumented
GoogleOther-ImageGoogleOther-ImageGoogledocumented
GoogleOther-VideoGoogleOther-VideoGoogledocumented
GPTBotGPTBotOpenAIdocumented
ICC-CrawlerICC-CrawlerNICTdocumented
ISSCyberRiskCrawlerISSCyberRiskCrawlerISS Corporate Solutionsdisputed
Linguee BotLinguee BotLingueedisputed
meta-externalagentmeta-externalagentMetadocumented
Poseidon Research CrawlerPoseidon Research CrawlerPoseidon Researchundocumented
QuillBotQuillBotQuillBotundocumented
ReflectionbotReflectionbotReflection AIundocumented
SBIntuitionsBotSBIntuitionsBotSB Intuitionsdocumented
SemrushBot-OCOBSemrushBot-OCOBSemrushdocumented
Sidetrade indexer botSidetrade indexer botSidetradeundocumented
TikTokSpiderTikTokSpiderByteDancedisputed
Webzio-ExtendedWebzio-ExtendedWebz.iodocumented
YandexAdditionalYandexAdditionalYandexown-token-only
YandexAdditionalBotYandexAdditionalBotYandexown-token-only

Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space.

Crawlerrobots.txt tokenOperatorrobots.txt
AIWebIndexAIWebIndexLyrenthdocumented
AmazonbotAmazonbotAmazondocumented
AndibotAndibotAndiundocumented
AnomuraAnomuraDireqtdocumented
atlassian-botatlassian-botAtlassiandocumented
bedrockbotbedrockbotAmazondocumented
Claude-SearchBotClaude-SearchBotAnthropicdocumented
Claude-WebClaude-WebAnthropicn-a
Cloudflare-AutoRAGCloudflare-AutoRAGCloudflaredocumented
DuckAssistBotDuckAssistBotDuckDuckGodocumented
ExaSearchBotExaSearchBotExaundocumented
Google-CloudVertexBotGoogle-CloudVertexBotGoogledocumented
KlaviyoAIBotKlaviyoAIBotKlaviyodocumented
Meta-WebIndexerMeta-WebIndexerMetaundocumented
OAI-SearchBotOAI-SearchBotOpenAIdocumented
PerplexityBotPerplexityBotPerplexitydocumented
PhindBotPhindBotPhindundocumented
QualifiedBotQualifiedBotQualifiedundocumented
ShapBotShapBotParalleldocumented
TerraCottaTerraCottaCeramic AIdocumented
YouBotYouBotYou.comdocumented

User-triggered fetchers (12)

Fetch one page because a person asked for it, right then. One human intent, one request. Blocking them produces a visible error for a real reader.

Crawlerrobots.txt tokenOperatorrobots.txt
ChatGPT AgentChatGPT-UserOpenAIdocumented
ChatGPT-UserChatGPT-UserOpenAIdocumented
Claude-UserClaude-UserAnthropicdocumented
cohere-aicohere-aiCoheredocumented
Google-AgentGoogle-AgentGoogleby-design-no
Google-GeminiNotebookGoogle-GeminiNotebookGoogleby-design-no
Google-PinpointGoogle-PinpointGoogleby-design-no
Google-Read-AloudGoogle-Read-AloudGoogleby-design-no
meta-externalfetchermeta-externalfetcherMetadocumented
MistralAI-UserMistralAI-UserMistral AIdocumented
Perplexity-UserPerplexity-UserPerplexityby-design-no
YandexCalendarYandexCalendarYandexown-token-only

Corpus and dataset builders (17)

Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect.

Crawlerrobots.txt tokenOperatorrobots.txt
AI2BotAI2BotAllen Institute for AIdocumented
Ai2Bot-DolmaAi2Bot-DolmaAllen Institute for AIdocumented
aiHitBotaiHitBotaiHitdocumented
AwarioRssBotAwarioRssBotAwariodocumented
AwarioSmartBotAwarioSmartBotAwariodocumented
CCBotCCBotCommon Crawldocumented
DiffbotDiffbotDiffbotdocumented
EchoboxBotEchoboxBotEchoboxundocumented
ImagesiftBotImagesiftBotHive AIdocumented
img2datasetimg2datasetLAION / img2datasetdocumented
LAIONDownloaderLAIONDownloaderLAION / img2datasetby-design-no
omgiliomgiliWebz.iodocumented
omgilibotomgilibotWebz.iodocumented
Panscientpanscient.comPanscientdocumented
ThinkbotThinkbotThinkbotdisputed
VelenPublicWebCrawlerVelenPublicWebCrawlerHunter (Velen)documented
YaKYaKMeltwaterundocumented

Classic index-and-rank crawlers. Several also feed their operator's generative answers, which is why the AI opt-out for Google and Apple is a token rather than a block.

Crawlerrobots.txt tokenOperatorrobots.txt
ApplebotApplebotAppledocumented
BaiduspiderBaiduspiderBaidudocumented
bingbotbingbotMicrosoftdocumented
DuckDuckBotDuckDuckBotDuckDuckGodocumented
GooglebotGooglebotGoogledocumented
Googlebot-ImageGooglebot-ImageGoogledocumented
Googlebot-NewsGooglebot-NewsGoogledocumented
Googlebot-VideoGooglebot-VideoGoogledocumented
KagibotKagibotKagidocumented
MojeekBotMojeekBotMojeekdocumented
PetalBotPetalBotHuaweidocumented
PinterestbotPinterestbotPinterestdocumented
QwantbotQwantbotQwantdocumented
Qwantbot-newsQwantbot-newsQwantdocumented
SeznamBotSeznamBotSeznamdocumented
Storebot-GoogleStorebot-GoogleGoogledocumented
TimpibotTimpibotTimpidocumented
YandexBlogsYandexBlogsYandexdocumented
YandexBotYandexBotYandexdocumented
YandexComBotYandexComBotYandexown-token-only
YandexFaviconsYandexFaviconsYandexown-token-only
YandexImagesYandexImagesYandexdocumented
YandexMarketYandexMarketYandexdocumented
YandexMediaYandexMediaYandexdocumented
YandexMobileBotYandexMobileBotYandexown-token-only
YandexRenderResourcesBotYandexRenderResourcesBotYandexown-token-only
YandexVideoYandexVideoYandexdocumented
YetiYetiNaverdocumented

SEO and backlink crawlers (17)

Commercial link-graph tooling. No user-facing effect either way, and usually a large share of your bot bandwidth.

Crawlerrobots.txt tokenOperatorrobots.txt
AhrefsBotAhrefsBotAhrefsdocumented
AhrefsSiteAuditAhrefsSiteAuditAhrefsdocumented
BarkrowlerbarkrowlerBabbardocumented
DataForSeoBotDataForSeoBotDataForSEOdocumented
DotBotdotbotMozdocumented
MJ12botMJ12botMajesticdocumented
rogerbotrogerbotMozdocumented
SemrushBotSemrushBotSemrushdocumented
SemrushBot-BASemrushBot-BASemrushdocumented
SemrushBot-ESISemrushBot-ESISemrushdocumented
SemrushBot-FTSemrushBot-FTSemrushdocumented
SemrushBot-SISemrushBot-SISemrushdocumented
SemrushBot-SWASemrushBot-SWASemrushdocumented
SEOkicksSEOkicksSEOkicksdocumented
serpstatbotserpstatbotSerpstatdocumented
SiteAuditBotSiteAuditBotSemrushdocumented
SplitSignalBotSplitSignalBotSemrushdocumented

Archivers (2)

Preservation crawlers. Their output is public and permanent, which makes them a separate decision from the AI one.

Crawlerrobots.txt tokenOperatorrobots.txt
archive.org_botarchive.org_botInternet Archivedocumented
ia_archiveria_archiverInternet Archivedocumented

Link preview fetchers (4)

Read your Open Graph tags when someone shares a link. Blocking these is almost always an accident.

Crawlerrobots.txt tokenOperatorrobots.txt
facebookexternalhitfacebookexternalhitMetadocumented
GoogleMessagesGoogleMessagesGoogleby-design-no
SlackbotSlackbotSlackdocumented
Slackbot-LinkExpandingSlackbot-LinkExpandingSlackdocumented

Tools and frameworks (22)

Not operators: crawling software anyone can run. The party behind the request is unknown, so treat them as a rate-limit question rather than a consent question.

Crawlerrobots.txt tokenOperatorrobots.txt
AdsBot-GoogleAdsBot-GoogleGoogleown-token-only
AdsBot-Google-MobileAdsBot-Google-MobileGoogleown-token-only
AdsBot-Google-Mobile-AppsAdsBot-Google-Mobile-AppsGoogleown-token-only
APIs-GoogleAPIs-GoogleGoogleown-token-only
Crawl4AICrawl4AICrawl4AI projectundocumented
CrawlspaceCrawlspaceCrawlspacedocumented
FeedFetcher-GoogleFeedFetcher-GoogleGoogleby-design-no
FirecrawlAgentFirecrawlAgentFirecrawldocumented
Google-CWSGoogle-CWSGoogleby-design-no
Google-InspectionToolGoogle-InspectionToolGoogledocumented
Google-SafetyGoogle-SafetyGoogleby-design-no
Google-Site-VerificationGoogle-Site-VerificationGoogleby-design-no
GoogleProducerGoogleProducerGoogleby-design-no
LightpandaLightpandaLightpandaundocumented
Mediapartners-GoogleMediapartners-GoogleGoogleown-token-only
ScrapyScrapyScrapy projectdocumented
Screaming Frog SEO SpiderScreaming Frog SEO SpiderScreaming Frogdocumented
wpbotwpbotQuantumCloudundocumented
YandexDirectYandexDirectYandexown-token-only
YandexMetrikaYandexMetrikaYandexby-design-no
YandexScreenshotBotYandexScreenshotBotYandexown-token-only
YandexWebmasterYandexWebmasterYandexdocumented