curl -s https://www.pathwren.workers.dev/policy/allow-ai-search-only.json   # this page, as JSON

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

Allow AI search and user fetches, block the rest

Be findable and citable in assistants without contributing to training corpora.

curl -s https://www.pathwren.workers.dev/robots/allow-ai-search-only.txt >> robots.txt

The inverse framing of block-ai-training, written as an allowlist so the default for anything new is deny. Fetches a user explicitly asked for stay allowed, because refusing those produces a visible error for a real person who wanted your page.

Names 65 crawlers

AIWebIndex · Amazonbot · Andibot · Anomura · Applebot · atlassian-bot · Baiduspider · bedrockbot · bingbot · ChatGPT Agent · ChatGPT-User · Claude-SearchBot · Claude-User · Claude-Web · Cloudflare-AutoRAG · cohere-ai · DuckAssistBot · DuckDuckBot · ExaSearchBot · facebookexternalhit · Google-Agent · Google-CloudVertexBot · Google-GeminiNotebook · Google-Pinpoint · Google-Read-Aloud · Googlebot · Googlebot-Image · Googlebot-News · Googlebot-Video · GoogleMessages · Kagibot · KlaviyoAIBot · meta-externalfetcher · Meta-WebIndexer · MistralAI-User · MojeekBot · OAI-SearchBot · Perplexity-User · PerplexityBot · PetalBot · PhindBot · Pinterestbot · QualifiedBot · Qwantbot · Qwantbot-news · SeznamBot · ShapBot · Slackbot · Slackbot-LinkExpanding · Storebot-Google · TerraCotta · Timpibot · YandexBlogs · YandexBot · YandexCalendar · YandexComBot · YandexFavicons · YandexImages · YandexMarket · YandexMedia · YandexMobileBot · YandexRenderResourcesBot · YandexVideo · Yeti · YouBot

The file

/robots/allow-ai-search-only.txt · json

# AI Crawler Index — policy: allow-ai-search-only
# Allow AI search and user fetches, block the rest
# Be findable and citable in assistants without contributing to training corpora.
# Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/allow-ai-search-only.html
# 65 crawlers named. Paste into robots.txt at your document root.

User-agent: AIWebIndex
Allow: /

User-agent: Amazonbot
Allow: /

User-agent: Andibot
Allow: /

User-agent: Anomura
Allow: /

User-agent: Applebot
Allow: /

User-agent: atlassian-bot
Allow: /

User-agent: Baiduspider
Allow: /

User-agent: bedrockbot
Allow: /

User-agent: bingbot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Claude-Web   # control token, no crawler uses this user-agent
Allow: /

User-agent: Cloudflare-AutoRAG
Allow: /

User-agent: cohere-ai
Allow: /

User-agent: DuckAssistBot
Allow: /

User-agent: DuckDuckBot
Allow: /

User-agent: ExaSearchBot
Allow: /

User-agent: facebookexternalhit
Allow: /

User-agent: Google-Agent   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-CloudVertexBot
Allow: /

User-agent: Google-GeminiNotebook   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-Pinpoint   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-Read-Aloud   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Googlebot
Allow: /

User-agent: Googlebot-Image
Allow: /

User-agent: Googlebot-News
Allow: /

User-agent: Googlebot-Video
Allow: /

User-agent: GoogleMessages   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Kagibot
Allow: /

User-agent: KlaviyoAIBot
Allow: /

User-agent: meta-externalfetcher
Allow: /

User-agent: Meta-WebIndexer
Allow: /

User-agent: MistralAI-User
Allow: /

User-agent: MojeekBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Perplexity-User   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: PetalBot
Allow: /

User-agent: PhindBot
Allow: /

User-agent: Pinterestbot
Allow: /

User-agent: QualifiedBot
Allow: /

User-agent: Qwantbot
Allow: /

User-agent: Qwantbot-news
Allow: /

User-agent: SeznamBot
Allow: /

User-agent: ShapBot
Allow: /

User-agent: Slackbot
Allow: /

User-agent: Slackbot-LinkExpanding
Allow: /

User-agent: Storebot-Google
Allow: /

User-agent: TerraCotta
Allow: /

User-agent: Timpibot
Allow: /

User-agent: YandexBlogs
Allow: /

User-agent: YandexBot
Allow: /

User-agent: YandexCalendar
Allow: /

User-agent: YandexComBot
Allow: /

User-agent: YandexFavicons
Allow: /

User-agent: YandexImages
Allow: /

User-agent: YandexMarket
Allow: /

User-agent: YandexMedia
Allow: /

User-agent: YandexMobileBot
Allow: /

User-agent: YandexRenderResourcesBot
Allow: /

User-agent: YandexVideo
Allow: /

User-agent: Yeti
Allow: /

User-agent: YouBot
Allow: /

# Anything not named above is refused.
User-agent: *
Disallow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml