curl -s https://www.pathwren.workers.dev/category/ai-training.json   # this page, as JSON

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

AI training crawlers

Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today.

CrawlerTokenOperatorCost of blocking
anthropic-aianthropic-aiAnthropicNone. Nothing crawls under this name today; keeping the rule is harmless insurance.…
Applebot-ExtendedApplebot-ExtendedAppleExcluded from Apple Intelligence training. Siri, Spotlight and Safari suggestions are unaf…
BytespiderBytespiderByteDanceLittle to lose. If you want it gone, expect to block by user-agent at the edge rather than…
ClaudeBotClaudeBotAnthropicContent excluded from training data for future Claude models. No effect on Claude's abilit…
cohere-training-data-crawlercohere-training-data-crawlerCohereExcluded from Cohere model training.…
CotoyogiCotoyogiROIS-DSYour Japanese-language content is left out of an academic training corpus.…
FacebookBotFacebookBotMetaNegligible today. Keep the rule; expect little traffic.…
Factset_spyderbotFactset_spyderbotFactSetExclusion from a financial-data vendor's corpus. Relevant mostly to companies whose filing…
Google-ExtendedGoogle-ExtendedGoogleYou are excluded from Gemini grounding and Gemini training. Google Search ranking and inde…
GoogleOtherGoogleOtherGoogleNo effect on Search indexing. Blocks internal Google research and product fetches.…
GoogleOther-ImageGoogleOther-ImageGoogleGoogle teams outside Search stop fetching your images. Image Search itself is unaffected —…
GoogleOther-VideoGoogleOther-VideoGoogleNo effect on Search or on Google Video search. Blocks internal Google research fetches of …
GPTBotGPTBotOpenAIYour content is excluded from training data for future OpenAI models. No effect on ChatGPT…
ICC-CrawlerICC-CrawlerNICTYou are excluded from a national research corpus and from the commercial redistributions o…
ISSCyberRiskCrawlerISSCyberRiskCrawlerISS Corporate SolutionsA rule here is a statement of intent. Your organisation's public footprint still gets scor…
Linguee BotLinguee BotLingueeMultilingual pages stop feeding a translation corpus. If your site is translated, being in…
meta-externalagentmeta-externalagentMetaExcluded from Meta AI training. Link previews on Facebook, Instagram and WhatsApp are unaf…
Poseidon Research CrawlerPoseidon Research CrawlerPoseidon ResearchExclusion from an interpretability research corpus. No published compliance statement.…
QuillBotQuillBotQuillBotExclusion from QuillBot's corpus. No compliance statement is published, so the rule is a r…
ReflectionbotReflectionbotReflection AIUnknown by construction — which is itself the reason some people block it. Nothing user-fa…
SBIntuitionsBotSBIntuitionsBotSB IntuitionsYour content is excluded from a Japanese-language foundation-model corpus. Nothing user-fa…
SemrushBot-OCOBSemrushBot-OCOBSemrushExclusion from Semrush's AI corpus, with its SEO crawl unaffected.…
Sidetrade indexer botSidetrade indexer botSidetradeExclusion from a commercial B2B dataset. The operator publishes no robots.txt statement, s…
TikTokSpiderTikTokSpiderByteDanceLittle to lose unless TikTok search referral matters to you.…
Webzio-ExtendedWebzio-ExtendedWebz.ioYour content is excluded from the AI-training tier of Webz.io's product while ordinary col…
YandexAdditionalYandexAdditionalYandexYou disappear from Yandex's AI answers while staying in Yandex Search. This is Yandex's eq…
YandexAdditionalBotYandexAdditionalBotYandexSame as YandexAdditional: out of Yandex's AI answers, still in Yandex Search. Name both to…

json