curl -s https://www.pathwren.workers.dev/policy/block-ai-training.json   # this page, as JSON

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

Block AI training, keep AI search

Refuse the crawlers that feed model training. Keep the ones that put you in ChatGPT, Claude, Perplexity and Gemini answers.

curl -s https://www.pathwren.workers.dev/robots/block-ai-training.txt >> robots.txt

The distinction most people actually want, and the one that is easy to get wrong: GPTBot trains, OAI-SearchBot indexes for citation. Blocking both loses you the traffic and gains you nothing extra. Google and Apple have no separate crawler at all — Google-Extended and Applebot-Extended are pure control tokens, so they belong in this file while Googlebot and Applebot must not.

Names 27 crawlers

anthropic-ai · Applebot-Extended · Bytespider · ClaudeBot · cohere-training-data-crawler · Cotoyogi · FacebookBot · Factset_spyderbot · Google-Extended · GoogleOther · GoogleOther-Image · GoogleOther-Video · GPTBot · ICC-Crawler · ISSCyberRiskCrawler · Linguee Bot · meta-externalagent · Poseidon Research Crawler · QuillBot · Reflectionbot · SBIntuitionsBot · SemrushBot-OCOB · Sidetrade indexer bot · TikTokSpider · Webzio-Extended · YandexAdditional · YandexAdditionalBot

The file

/robots/block-ai-training.txt · json

# AI Crawler Index — policy: block-ai-training
# Block AI training, keep AI search
# Refuse the crawlers that feed model training. Keep the ones that put you in ChatGPT, Claude, Perplexity and Gemini answers.
# Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/block-ai-training.html
# 27 crawlers named. Paste into robots.txt at your document root.

User-agent: anthropic-ai   # control token, no crawler uses this user-agent
Disallow: /

User-agent: Applebot-Extended   # control token, no crawler uses this user-agent
Disallow: /

User-agent: Bytespider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: cohere-training-data-crawler
Disallow: /

User-agent: Cotoyogi
Disallow: /

User-agent: FacebookBot
Disallow: /

User-agent: Factset_spyderbot
Disallow: /

User-agent: Google-Extended   # control token, no crawler uses this user-agent
Disallow: /

User-agent: GoogleOther
Disallow: /

User-agent: GoogleOther-Image
Disallow: /

User-agent: GoogleOther-Video
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: ICC-Crawler
Disallow: /

User-agent: ISSCyberRiskCrawler   # compliance disputed; enforce at the edge
Disallow: /

User-agent: Linguee Bot   # compliance disputed; enforce at the edge
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: Poseidon Research Crawler
Disallow: /

User-agent: QuillBot
Disallow: /

User-agent: Reflectionbot
Disallow: /

User-agent: SBIntuitionsBot
Disallow: /

User-agent: SemrushBot-OCOB
Disallow: /

User-agent: Sidetrade indexer bot
Disallow: /

User-agent: TikTokSpider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: Webzio-Extended
Disallow: /

User-agent: YandexAdditional
Disallow: /

User-agent: YandexAdditionalBot
Disallow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml