curl -s https://www.pathwren.workers.dev/policy/block-ai-training.json # this page, as JSON
No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a
Refuse the crawlers that feed model training. Keep the ones that put you in ChatGPT, Claude, Perplexity and Gemini answers.
curl -s https://www.pathwren.workers.dev/robots/block-ai-training.txt >> robots.txt
The distinction most people actually want, and the one that is easy to get wrong: GPTBot trains, OAI-SearchBot indexes for citation. Blocking both loses you the traffic and gains you nothing extra. Google and Apple have no separate crawler at all — Google-Extended and Applebot-Extended are pure control tokens, so they belong in this file while Googlebot and Applebot must not.
anthropic-ai · Applebot-Extended · Bytespider · ClaudeBot · cohere-training-data-crawler · Cotoyogi · FacebookBot · Factset_spyderbot · Google-Extended · GoogleOther · GoogleOther-Image · GoogleOther-Video · GPTBot · ICC-Crawler · ISSCyberRiskCrawler · Linguee Bot · meta-externalagent · Poseidon Research Crawler · QuillBot · Reflectionbot · SBIntuitionsBot · SemrushBot-OCOB · Sidetrade indexer bot · TikTokSpider · Webzio-Extended · YandexAdditional · YandexAdditionalBot
/robots/block-ai-training.txt · json
# AI Crawler Index — policy: block-ai-training # Block AI training, keep AI search # Refuse the crawlers that feed model training. Keep the ones that put you in ChatGPT, Claude, Perplexity and Gemini answers. # Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/block-ai-training.html # 27 crawlers named. Paste into robots.txt at your document root. User-agent: anthropic-ai # control token, no crawler uses this user-agent Disallow: / User-agent: Applebot-Extended # control token, no crawler uses this user-agent Disallow: / User-agent: Bytespider # compliance disputed; enforce at the edge Disallow: / User-agent: ClaudeBot Disallow: / User-agent: cohere-training-data-crawler Disallow: / User-agent: Cotoyogi Disallow: / User-agent: FacebookBot Disallow: / User-agent: Factset_spyderbot Disallow: / User-agent: Google-Extended # control token, no crawler uses this user-agent Disallow: / User-agent: GoogleOther Disallow: / User-agent: GoogleOther-Image Disallow: / User-agent: GoogleOther-Video Disallow: / User-agent: GPTBot Disallow: / User-agent: ICC-Crawler Disallow: / User-agent: ISSCyberRiskCrawler # compliance disputed; enforce at the edge Disallow: / User-agent: Linguee Bot # compliance disputed; enforce at the edge Disallow: / User-agent: meta-externalagent Disallow: / User-agent: Poseidon Research Crawler Disallow: / User-agent: QuillBot Disallow: / User-agent: Reflectionbot Disallow: / User-agent: SBIntuitionsBot Disallow: / User-agent: SemrushBot-OCOB Disallow: / User-agent: Sidetrade indexer bot Disallow: / User-agent: TikTokSpider # compliance disputed; enforce at the edge Disallow: / User-agent: Webzio-Extended Disallow: / User-agent: YandexAdditional Disallow: / User-agent: YandexAdditionalBot Disallow: / User-agent: * Allow: / Sitemap: https://www.pathwren.workers.dev/sitemap.xml