---
title: "robots.txt: Block every AI crawler — AI Crawler Index"
description: "Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed."
canonical: "https://www.pathwren.workers.dev/policy/block-all-ai.html"
url: "https://www.pathwren.workers.dev/policy/block-all-ai.md"
format: "markdown"
source: "the bytes of /policy/block-all-ai.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T06:29:54+00:00"
license: "CC0-1.0"
---

# Block every AI crawler

> Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed.

Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed.

```bash
curl -s https://www.pathwren.workers.dev/robots/block-all-ai.txt >> robots.txt
```

The maximal AI opt-out that still leaves you in Google and Bing. Understand the price before deploying it: you will not be cited by any assistant, and when a reader explicitly asks ChatGPT or Claude to open your page, they get an error. Note also that Perplexity-User and Bytespider are listed here but documented as not governed by robots.txt, so this file is a statement of intent for those two, not an enforcement mechanism.

## Names 77 crawlers

[AI2Bot](https://www.pathwren.workers.dev/crawler/ai2bot.html) · [Ai2Bot-Dolma](https://www.pathwren.workers.dev/crawler/ai2bot-dolma.html) · [aiHitBot](https://www.pathwren.workers.dev/crawler/aihitbot.html) · [AIWebIndex](https://www.pathwren.workers.dev/crawler/aiwebindex.html) · [Amazonbot](https://www.pathwren.workers.dev/crawler/amazonbot.html) · [Andibot](https://www.pathwren.workers.dev/crawler/andibot.html) · [Anomura](https://www.pathwren.workers.dev/crawler/anomura.html) · [anthropic-ai](https://www.pathwren.workers.dev/crawler/anthropic-ai.html) · [Applebot-Extended](https://www.pathwren.workers.dev/crawler/applebot-extended.html) · [atlassian-bot](https://www.pathwren.workers.dev/crawler/atlassian-bot.html) · [AwarioRssBot](https://www.pathwren.workers.dev/crawler/awariorssbot.html) · [AwarioSmartBot](https://www.pathwren.workers.dev/crawler/awariosmartbot.html) · [bedrockbot](https://www.pathwren.workers.dev/crawler/bedrockbot.html) · [Bytespider](https://www.pathwren.workers.dev/crawler/bytespider.html) · [CCBot](https://www.pathwren.workers.dev/crawler/ccbot.html) · [ChatGPT Agent](https://www.pathwren.workers.dev/crawler/chatgpt-agent.html) · [ChatGPT-User](https://www.pathwren.workers.dev/crawler/chatgpt-user.html) · [Claude-SearchBot](https://www.pathwren.workers.dev/crawler/claude-searchbot.html) · [Claude-User](https://www.pathwren.workers.dev/crawler/claude-user.html) · [Claude-Web](https://www.pathwren.workers.dev/crawler/claude-web.html) · [ClaudeBot](https://www.pathwren.workers.dev/crawler/claudebot.html) · [Cloudflare-AutoRAG](https://www.pathwren.workers.dev/crawler/cloudflare-autorag.html) · [cohere-ai](https://www.pathwren.workers.dev/crawler/cohere-ai.html) · [cohere-training-data-crawler](https://www.pathwren.workers.dev/crawler/cohere-training-data-crawler.html) · [Cotoyogi](https://www.pathwren.workers.dev/crawler/cotoyogi.html) · [Diffbot](https://www.pathwren.workers.dev/crawler/diffbot.html) · [DuckAssistBot](https://www.pathwren.workers.dev/crawler/duckassistbot.html) · [EchoboxBot](https://www.pathwren.workers.dev/crawler/echoboxbot.html) · [ExaSearchBot](https://www.pathwren.workers.dev/crawler/exasearchbot.html) · [FacebookBot](https://www.pathwren.workers.dev/crawler/facebookbot.html) · [Factset_spyderbot](https://www.pathwren.workers.dev/crawler/factset-spyderbot.html) · [Google-Agent](https://www.pathwren.workers.dev/crawler/google-agent.html) · [Google-CloudVertexBot](https://www.pathwren.workers.dev/crawler/google-cloudvertexbot.html) · [Google-Extended](https://www.pathwren.workers.dev/crawler/google-extended.html) · [Google-GeminiNotebook](https://www.pathwren.workers.dev/crawler/google-gemininotebook.html) · [Google-Pinpoint](https://www.pathwren.workers.dev/crawler/google-pinpoint.html) · [Google-Read-Aloud](https://www.pathwren.workers.dev/crawler/google-read-aloud.html) · [GoogleOther](https://www.pathwren.workers.dev/crawler/googleother.html) · [GoogleOther-Image](https://www.pathwren.workers.dev/crawler/googleother-image.html) · [GoogleOther-Video](https://www.pathwren.workers.dev/crawler/googleother-video.html) · [GPTBot](https://www.pathwren.workers.dev/crawler/gptbot.html) · [ICC-Crawler](https://www.pathwren.workers.dev/crawler/icc-crawler.html) · [ImagesiftBot](https://www.pathwren.workers.dev/crawler/imagesiftbot.html) · [img2dataset](https://www.pathwren.workers.dev/crawler/img2dataset.html) · [ISSCyberRiskCrawler](https://www.pathwren.workers.dev/crawler/isscyberriskcrawler.html) · [KlaviyoAIBot](https://www.pathwren.workers.dev/crawler/klaviyoaibot.html) · [LAIONDownloader](https://www.pathwren.workers.dev/crawler/laiondownloader.html) · [Linguee Bot](https://www.pathwren.workers.dev/crawler/linguee-bot.html) · [meta-externalagent](https://www.pathwren.workers.dev/crawler/meta-externalagent.html) · [meta-externalfetcher](https://www.pathwren.workers.dev/crawler/meta-externalfetcher.html) · [Meta-WebIndexer](https://www.pathwren.workers.dev/crawler/meta-webindexer.html) · [MistralAI-User](https://www.pathwren.workers.dev/crawler/mistralai-user.html) · [OAI-SearchBot](https://www.pathwren.workers.dev/crawler/oai-searchbot.html) · [omgili](https://www.pathwren.workers.dev/crawler/omgili.html) · [omgilibot](https://www.pathwren.workers.dev/crawler/omgilibot.html) · [Panscient](https://www.pathwren.workers.dev/crawler/panscient.html) · [Perplexity-User](https://www.pathwren.workers.dev/crawler/perplexity-user.html) · [PerplexityBot](https://www.pathwren.workers.dev/crawler/perplexitybot.html) · [PhindBot](https://www.pathwren.workers.dev/crawler/phindbot.html) · [Poseidon Research Crawler](https://www.pathwren.workers.dev/crawler/poseidon-research-crawler.html) · [QualifiedBot](https://www.pathwren.workers.dev/crawler/qualifiedbot.html) · [QuillBot](https://www.pathwren.workers.dev/crawler/quillbot.html) · [Reflectionbot](https://www.pathwren.workers.dev/crawler/reflectionbot.html) · [SBIntuitionsBot](https://www.pathwren.workers.dev/crawler/sbintuitionsbot.html) · [SemrushBot-OCOB](https://www.pathwren.workers.dev/crawler/semrushbot-ocob.html) · [ShapBot](https://www.pathwren.workers.dev/crawler/shapbot.html) · [Sidetrade indexer bot](https://www.pathwren.workers.dev/crawler/sidetrade-indexer-bot.html) · [TerraCotta](https://www.pathwren.workers.dev/crawler/terracotta.html) · [Thinkbot](https://www.pathwren.workers.dev/crawler/thinkbot.html) · [TikTokSpider](https://www.pathwren.workers.dev/crawler/tiktokspider.html) · [VelenPublicWebCrawler](https://www.pathwren.workers.dev/crawler/velenpublicwebcrawler.html) · [Webzio-Extended](https://www.pathwren.workers.dev/crawler/webzio-extended.html) · [YaK](https://www.pathwren.workers.dev/crawler/yak.html) · [YandexAdditional](https://www.pathwren.workers.dev/crawler/yandexadditional.html) · [YandexAdditionalBot](https://www.pathwren.workers.dev/crawler/yandexadditionalbot.html) · [YandexCalendar](https://www.pathwren.workers.dev/crawler/yandexcalendar.html) · [YouBot](https://www.pathwren.workers.dev/crawler/youbot.html)

## The file

[/robots/block-all-ai.txt](https://www.pathwren.workers.dev/robots/block-all-ai.txt) · [json](https://www.pathwren.workers.dev/policy/block-all-ai.json)

```text
# AI Crawler Index — policy: block-all-ai
# Block every AI crawler
# Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed.
# Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/block-all-ai.html
# 77 crawlers named. Paste into robots.txt at your document root.

User-agent: AI2Bot
Disallow: /

User-agent: Ai2Bot-Dolma
Disallow: /

User-agent: aiHitBot
Disallow: /

User-agent: AIWebIndex
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Andibot
Disallow: /

User-agent: Anomura
Disallow: /

User-agent: anthropic-ai   # control token, no crawler uses this user-agent
Disallow: /

User-agent: Applebot-Extended   # control token, no crawler uses this user-agent
Disallow: /

User-agent: atlassian-bot
Disallow: /

User-agent: AwarioRssBot
Disallow: /

User-agent: AwarioSmartBot
Disallow: /

User-agent: bedrockbot
Disallow: /

User-agent: Bytespider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: Claude-Web   # control token, no crawler uses this user-agent
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Cloudflare-AutoRAG
Disallow: /

User-agent: cohere-ai
Disallow: /

User-agent: cohere-training-data-crawler
Disallow: /

User-agent: Cotoyogi
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: DuckAssistBot
Disallow: /

User-agent: EchoboxBot
Disallow: /

User-agent: ExaSearchBot
Disallow: /

User-agent: FacebookBot
Disallow: /

User-agent: Factset_spyderbot
Disallow: /

User-agent: Google-Agent   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-CloudVertexBot
Disallow: /

User-agent: Google-Extended   # control token, no crawler uses this user-agent
Disallow: /

User-agent: Google-GeminiNotebook   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Pinpoint   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Read-Aloud   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: GoogleOther
Disallow: /

User-agent: GoogleOther-Image
Disallow: /

User-agent: GoogleOther-Video
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: ICC-Crawler
Disallow: /

User-agent: ImagesiftBot
Disallow: /

User-agent: img2dataset
Disallow: /

User-agent: ISSCyberRiskCrawler   # compliance disputed; enforce at the edge
Disallow: /

User-agent: KlaviyoAIBot
Disallow: /

User-agent: LAIONDownloader   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Linguee Bot   # compliance disputed; enforce at the edge
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: meta-externalfetcher
Disallow: /

User-agent: Meta-WebIndexer
Disallow: /

User-agent: MistralAI-User
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: omgili
Disallow: /

User-agent: omgilibot
Disallow: /

User-agent: panscient.com
Disallow: /

User-agent: Perplexity-User   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: PhindBot
Disallow: /

User-agent: Poseidon Research Crawler
Disallow: /

User-agent: QualifiedBot
Disallow: /

User-agent: QuillBot
Disallow: /

User-agent: Reflectionbot
Disallow: /

User-agent: SBIntuitionsBot
Disallow: /

User-agent: SemrushBot-OCOB
Disallow: /

User-agent: ShapBot
Disallow: /

User-agent: Sidetrade indexer bot
Disallow: /

User-agent: TerraCotta
Disallow: /

User-agent: Thinkbot   # compliance disputed; enforce at the edge
Disallow: /

User-agent: TikTokSpider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: VelenPublicWebCrawler
Disallow: /

User-agent: Webzio-Extended
Disallow: /

User-agent: YaK
Disallow: /

User-agent: YandexAdditional
Disallow: /

User-agent: YandexAdditionalBot
Disallow: /

User-agent: YandexCalendar
Disallow: /

User-agent: YouBot
Disallow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml
```

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/policy/block-all-ai.html)
- [JSON](https://www.pathwren.workers.dev/policy/block-all-ai.json)
- [Markdown](https://www.pathwren.workers.dev/policy/block-all-ai.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/policy/block-all-ai.html](https://www.pathwren.workers.dev/policy/block-all-ai.html), generated from that page's own bytes in the same build. The HTML page is canonical.
