---
title: "robots.txt: Allow AI search and user fetches, block the rest — AI Crawler Index"
description: "Be findable and citable in assistants without contributing to training corpora."
canonical: "https://www.pathwren.workers.dev/policy/allow-ai-search-only.html"
url: "https://www.pathwren.workers.dev/policy/allow-ai-search-only.md"
format: "markdown"
source: "the bytes of /policy/allow-ai-search-only.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T07:55:35+00:00"
license: "CC0-1.0"
---

# Allow AI search and user fetches, block the rest

> Be findable and citable in assistants without contributing to training corpora.

Be findable and citable in assistants without contributing to training corpora.

```bash
curl -s https://www.pathwren.workers.dev/robots/allow-ai-search-only.txt >> robots.txt
```

The inverse framing of block-ai-training, written as an allowlist so the default for anything new is deny. Fetches a user explicitly asked for stay allowed, because refusing those produces a visible error for a real person who wanted your page.

## Names 65 crawlers

[AIWebIndex](https://www.pathwren.workers.dev/crawler/aiwebindex.html) · [Amazonbot](https://www.pathwren.workers.dev/crawler/amazonbot.html) · [Andibot](https://www.pathwren.workers.dev/crawler/andibot.html) · [Anomura](https://www.pathwren.workers.dev/crawler/anomura.html) · [Applebot](https://www.pathwren.workers.dev/crawler/applebot.html) · [atlassian-bot](https://www.pathwren.workers.dev/crawler/atlassian-bot.html) · [Baiduspider](https://www.pathwren.workers.dev/crawler/baiduspider.html) · [bedrockbot](https://www.pathwren.workers.dev/crawler/bedrockbot.html) · [bingbot](https://www.pathwren.workers.dev/crawler/bingbot.html) · [ChatGPT Agent](https://www.pathwren.workers.dev/crawler/chatgpt-agent.html) · [ChatGPT-User](https://www.pathwren.workers.dev/crawler/chatgpt-user.html) · [Claude-SearchBot](https://www.pathwren.workers.dev/crawler/claude-searchbot.html) · [Claude-User](https://www.pathwren.workers.dev/crawler/claude-user.html) · [Claude-Web](https://www.pathwren.workers.dev/crawler/claude-web.html) · [Cloudflare-AutoRAG](https://www.pathwren.workers.dev/crawler/cloudflare-autorag.html) · [cohere-ai](https://www.pathwren.workers.dev/crawler/cohere-ai.html) · [DuckAssistBot](https://www.pathwren.workers.dev/crawler/duckassistbot.html) · [DuckDuckBot](https://www.pathwren.workers.dev/crawler/duckduckbot.html) · [ExaSearchBot](https://www.pathwren.workers.dev/crawler/exasearchbot.html) · [facebookexternalhit](https://www.pathwren.workers.dev/crawler/facebookexternalhit.html) · [Google-Agent](https://www.pathwren.workers.dev/crawler/google-agent.html) · [Google-CloudVertexBot](https://www.pathwren.workers.dev/crawler/google-cloudvertexbot.html) · [Google-GeminiNotebook](https://www.pathwren.workers.dev/crawler/google-gemininotebook.html) · [Google-Pinpoint](https://www.pathwren.workers.dev/crawler/google-pinpoint.html) · [Google-Read-Aloud](https://www.pathwren.workers.dev/crawler/google-read-aloud.html) · [Googlebot](https://www.pathwren.workers.dev/crawler/googlebot.html) · [Googlebot-Image](https://www.pathwren.workers.dev/crawler/googlebot-image.html) · [Googlebot-News](https://www.pathwren.workers.dev/crawler/googlebot-news.html) · [Googlebot-Video](https://www.pathwren.workers.dev/crawler/googlebot-video.html) · [GoogleMessages](https://www.pathwren.workers.dev/crawler/googlemessages.html) · [Kagibot](https://www.pathwren.workers.dev/crawler/kagibot.html) · [KlaviyoAIBot](https://www.pathwren.workers.dev/crawler/klaviyoaibot.html) · [meta-externalfetcher](https://www.pathwren.workers.dev/crawler/meta-externalfetcher.html) · [Meta-WebIndexer](https://www.pathwren.workers.dev/crawler/meta-webindexer.html) · [MistralAI-User](https://www.pathwren.workers.dev/crawler/mistralai-user.html) · [MojeekBot](https://www.pathwren.workers.dev/crawler/mojeekbot.html) · [OAI-SearchBot](https://www.pathwren.workers.dev/crawler/oai-searchbot.html) · [Perplexity-User](https://www.pathwren.workers.dev/crawler/perplexity-user.html) · [PerplexityBot](https://www.pathwren.workers.dev/crawler/perplexitybot.html) · [PetalBot](https://www.pathwren.workers.dev/crawler/petalbot.html) · [PhindBot](https://www.pathwren.workers.dev/crawler/phindbot.html) · [Pinterestbot](https://www.pathwren.workers.dev/crawler/pinterestbot.html) · [QualifiedBot](https://www.pathwren.workers.dev/crawler/qualifiedbot.html) · [Qwantbot](https://www.pathwren.workers.dev/crawler/qwantbot.html) · [Qwantbot-news](https://www.pathwren.workers.dev/crawler/qwantbot-news.html) · [SeznamBot](https://www.pathwren.workers.dev/crawler/seznambot.html) · [ShapBot](https://www.pathwren.workers.dev/crawler/shapbot.html) · [Slackbot](https://www.pathwren.workers.dev/crawler/slackbot.html) · [Slackbot-LinkExpanding](https://www.pathwren.workers.dev/crawler/slackbot-linkexpanding.html) · [Storebot-Google](https://www.pathwren.workers.dev/crawler/storebot-google.html) · [TerraCotta](https://www.pathwren.workers.dev/crawler/terracotta.html) · [Timpibot](https://www.pathwren.workers.dev/crawler/timpibot.html) · [YandexBlogs](https://www.pathwren.workers.dev/crawler/yandexblogs.html) · [YandexBot](https://www.pathwren.workers.dev/crawler/yandexbot.html) · [YandexCalendar](https://www.pathwren.workers.dev/crawler/yandexcalendar.html) · [YandexComBot](https://www.pathwren.workers.dev/crawler/yandexcombot.html) · [YandexFavicons](https://www.pathwren.workers.dev/crawler/yandexfavicons.html) · [YandexImages](https://www.pathwren.workers.dev/crawler/yandeximages.html) · [YandexMarket](https://www.pathwren.workers.dev/crawler/yandexmarket.html) · [YandexMedia](https://www.pathwren.workers.dev/crawler/yandexmedia.html) · [YandexMobileBot](https://www.pathwren.workers.dev/crawler/yandexmobilebot.html) · [YandexRenderResourcesBot](https://www.pathwren.workers.dev/crawler/yandexrenderresourcesbot.html) · [YandexVideo](https://www.pathwren.workers.dev/crawler/yandexvideo.html) · [Yeti](https://www.pathwren.workers.dev/crawler/yeti.html) · [YouBot](https://www.pathwren.workers.dev/crawler/youbot.html)

## The file

[/robots/allow-ai-search-only.txt](https://www.pathwren.workers.dev/robots/allow-ai-search-only.txt) · [json](https://www.pathwren.workers.dev/policy/allow-ai-search-only.json)

```text
# AI Crawler Index — policy: allow-ai-search-only
# Allow AI search and user fetches, block the rest
# Be findable and citable in assistants without contributing to training corpora.
# Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/allow-ai-search-only.html
# 65 crawlers named. Paste into robots.txt at your document root.

User-agent: AIWebIndex
Allow: /

User-agent: Amazonbot
Allow: /

User-agent: Andibot
Allow: /

User-agent: Anomura
Allow: /

User-agent: Applebot
Allow: /

User-agent: atlassian-bot
Allow: /

User-agent: Baiduspider
Allow: /

User-agent: bedrockbot
Allow: /

User-agent: bingbot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Claude-Web   # control token, no crawler uses this user-agent
Allow: /

User-agent: Cloudflare-AutoRAG
Allow: /

User-agent: cohere-ai
Allow: /

User-agent: DuckAssistBot
Allow: /

User-agent: DuckDuckBot
Allow: /

User-agent: ExaSearchBot
Allow: /

User-agent: facebookexternalhit
Allow: /

User-agent: Google-Agent   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-CloudVertexBot
Allow: /

User-agent: Google-GeminiNotebook   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-Pinpoint   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-Read-Aloud   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Googlebot
Allow: /

User-agent: Googlebot-Image
Allow: /

User-agent: Googlebot-News
Allow: /

User-agent: Googlebot-Video
Allow: /

User-agent: GoogleMessages   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Kagibot
Allow: /

User-agent: KlaviyoAIBot
Allow: /

User-agent: meta-externalfetcher
Allow: /

User-agent: Meta-WebIndexer
Allow: /

User-agent: MistralAI-User
Allow: /

User-agent: MojeekBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Perplexity-User   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: PetalBot
Allow: /

User-agent: PhindBot
Allow: /

User-agent: Pinterestbot
Allow: /

User-agent: QualifiedBot
Allow: /

User-agent: Qwantbot
Allow: /

User-agent: Qwantbot-news
Allow: /

User-agent: SeznamBot
Allow: /

User-agent: ShapBot
Allow: /

User-agent: Slackbot
Allow: /

User-agent: Slackbot-LinkExpanding
Allow: /

User-agent: Storebot-Google
Allow: /

User-agent: TerraCotta
Allow: /

User-agent: Timpibot
Allow: /

User-agent: YandexBlogs
Allow: /

User-agent: YandexBot
Allow: /

User-agent: YandexCalendar
Allow: /

User-agent: YandexComBot
Allow: /

User-agent: YandexFavicons
Allow: /

User-agent: YandexImages
Allow: /

User-agent: YandexMarket
Allow: /

User-agent: YandexMedia
Allow: /

User-agent: YandexMobileBot
Allow: /

User-agent: YandexRenderResourcesBot
Allow: /

User-agent: YandexVideo
Allow: /

User-agent: Yeti
Allow: /

User-agent: YouBot
Allow: /

# Anything not named above is refused.
User-agent: *
Disallow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml
```

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/policy/allow-ai-search-only.html)
- [JSON](https://www.pathwren.workers.dev/policy/allow-ai-search-only.json)
- [Markdown](https://www.pathwren.workers.dev/policy/allow-ai-search-only.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/policy/allow-ai-search-only.html](https://www.pathwren.workers.dev/policy/allow-ai-search-only.html), generated from that page's own bytes in the same build. The HTML page is canonical.
