curl -s https://www.pathwren.workers.dev/policy/block-all-ai.json   # this page, as JSON

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

Block every AI crawler

Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed.

curl -s https://www.pathwren.workers.dev/robots/block-all-ai.txt >> robots.txt

The maximal AI opt-out that still leaves you in Google and Bing. Understand the price before deploying it: you will not be cited by any assistant, and when a reader explicitly asks ChatGPT or Claude to open your page, they get an error. Note also that Perplexity-User and Bytespider are listed here but documented as not governed by robots.txt, so this file is a statement of intent for those two, not an enforcement mechanism.

Names 77 crawlers

AI2Bot · Ai2Bot-Dolma · aiHitBot · AIWebIndex · Amazonbot · Andibot · Anomura · anthropic-ai · Applebot-Extended · atlassian-bot · AwarioRssBot · AwarioSmartBot · bedrockbot · Bytespider · CCBot · ChatGPT Agent · ChatGPT-User · Claude-SearchBot · Claude-User · Claude-Web · ClaudeBot · Cloudflare-AutoRAG · cohere-ai · cohere-training-data-crawler · Cotoyogi · Diffbot · DuckAssistBot · EchoboxBot · ExaSearchBot · FacebookBot · Factset_spyderbot · Google-Agent · Google-CloudVertexBot · Google-Extended · Google-GeminiNotebook · Google-Pinpoint · Google-Read-Aloud · GoogleOther · GoogleOther-Image · GoogleOther-Video · GPTBot · ICC-Crawler · ImagesiftBot · img2dataset · ISSCyberRiskCrawler · KlaviyoAIBot · LAIONDownloader · Linguee Bot · meta-externalagent · meta-externalfetcher · Meta-WebIndexer · MistralAI-User · OAI-SearchBot · omgili · omgilibot · Panscient · Perplexity-User · PerplexityBot · PhindBot · Poseidon Research Crawler · QualifiedBot · QuillBot · Reflectionbot · SBIntuitionsBot · SemrushBot-OCOB · ShapBot · Sidetrade indexer bot · TerraCotta · Thinkbot · TikTokSpider · VelenPublicWebCrawler · Webzio-Extended · YaK · YandexAdditional · YandexAdditionalBot · YandexCalendar · YouBot

The file

/robots/block-all-ai.txt · json

# AI Crawler Index — policy: block-all-ai
# Block every AI crawler
# Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed.
# Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/block-all-ai.html
# 77 crawlers named. Paste into robots.txt at your document root.

User-agent: AI2Bot
Disallow: /

User-agent: Ai2Bot-Dolma
Disallow: /

User-agent: aiHitBot
Disallow: /

User-agent: AIWebIndex
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Andibot
Disallow: /

User-agent: Anomura
Disallow: /

User-agent: anthropic-ai   # control token, no crawler uses this user-agent
Disallow: /

User-agent: Applebot-Extended   # control token, no crawler uses this user-agent
Disallow: /

User-agent: atlassian-bot
Disallow: /

User-agent: AwarioRssBot
Disallow: /

User-agent: AwarioSmartBot
Disallow: /

User-agent: bedrockbot
Disallow: /

User-agent: Bytespider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: Claude-Web   # control token, no crawler uses this user-agent
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Cloudflare-AutoRAG
Disallow: /

User-agent: cohere-ai
Disallow: /

User-agent: cohere-training-data-crawler
Disallow: /

User-agent: Cotoyogi
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: DuckAssistBot
Disallow: /

User-agent: EchoboxBot
Disallow: /

User-agent: ExaSearchBot
Disallow: /

User-agent: FacebookBot
Disallow: /

User-agent: Factset_spyderbot
Disallow: /

User-agent: Google-Agent   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-CloudVertexBot
Disallow: /

User-agent: Google-Extended   # control token, no crawler uses this user-agent
Disallow: /

User-agent: Google-GeminiNotebook   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Pinpoint   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Read-Aloud   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: GoogleOther
Disallow: /

User-agent: GoogleOther-Image
Disallow: /

User-agent: GoogleOther-Video
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: ICC-Crawler
Disallow: /

User-agent: ImagesiftBot
Disallow: /

User-agent: img2dataset
Disallow: /

User-agent: ISSCyberRiskCrawler   # compliance disputed; enforce at the edge
Disallow: /

User-agent: KlaviyoAIBot
Disallow: /

User-agent: LAIONDownloader   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Linguee Bot   # compliance disputed; enforce at the edge
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: meta-externalfetcher
Disallow: /

User-agent: Meta-WebIndexer
Disallow: /

User-agent: MistralAI-User
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: omgili
Disallow: /

User-agent: omgilibot
Disallow: /

User-agent: panscient.com
Disallow: /

User-agent: Perplexity-User   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: PhindBot
Disallow: /

User-agent: Poseidon Research Crawler
Disallow: /

User-agent: QualifiedBot
Disallow: /

User-agent: QuillBot
Disallow: /

User-agent: Reflectionbot
Disallow: /

User-agent: SBIntuitionsBot
Disallow: /

User-agent: SemrushBot-OCOB
Disallow: /

User-agent: ShapBot
Disallow: /

User-agent: Sidetrade indexer bot
Disallow: /

User-agent: TerraCotta
Disallow: /

User-agent: Thinkbot   # compliance disputed; enforce at the edge
Disallow: /

User-agent: TikTokSpider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: VelenPublicWebCrawler
Disallow: /

User-agent: Webzio-Extended
Disallow: /

User-agent: YaK
Disallow: /

User-agent: YandexAdditional
Disallow: /

User-agent: YandexAdditionalBot
Disallow: /

User-agent: YandexCalendar
Disallow: /

User-agent: YouBot
Disallow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml