---
title: "AI search crawlers (21) — AI Crawler Index"
description: "Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space."
canonical: "https://www.pathwren.workers.dev/category/ai-search.html"
url: "https://www.pathwren.workers.dev/category/ai-search.md"
format: "markdown"
source: "the bytes of /category/ai-search.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T06:29:54+00:00"
license: "CC0-1.0"
---

# AI search crawlers

> Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space.

Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space.

| Crawler | Token | Operator | Cost of blocking |
| --- | --- | --- | --- |
| [AIWebIndex](https://www.pathwren.workers.dev/crawler/aiwebindex.html) | `AIWebIndex` | Lyrenth | Agents reading through this index stop seeing you — including the attribution and link bac… |
| [Amazonbot](https://www.pathwren.workers.dev/crawler/amazonbot.html) | `Amazonbot` | Amazon | Alexa and Amazon's assistants stop answering from your pages. Verify with reverse DNS to c… |
| [Andibot](https://www.pathwren.workers.dev/crawler/andibot.html) | `Andibot` | Andi | You disappear from another assistant's answers. Andi publishes no robots.txt statement.… |
| [Anomura](https://www.pathwren.workers.dev/crawler/anomura.html) | `Anomura` | Direqt | If you are the publisher, this breaks the assistant you put on your own pages. If you are … |
| [atlassian-bot](https://www.pathwren.workers.dev/crawler/atlassian-bot.html) | `atlassian-bot` | Atlassian | Rovo cannot answer from your public documentation. If your customers live inside Atlassian… |
| [bedrockbot](https://www.pathwren.workers.dev/crawler/bedrockbot.html) | `bedrockbot` | Amazon | Companies building retrieval applications on Bedrock cannot include your pages. This is a … |
| [Claude-SearchBot](https://www.pathwren.workers.dev/crawler/claude-searchbot.html) | `Claude-SearchBot` | Anthropic | You stop appearing in Claude's search results and citations.… |
| [Claude-Web](https://www.pathwren.workers.dev/crawler/claude-web.html) | `Claude-Web` | Anthropic | None in practice. Retain the rule; expect no traffic.… |
| [Cloudflare-AutoRAG](https://www.pathwren.workers.dev/crawler/cloudflare-autorag.html) | `Cloudflare-AutoRAG` | Cloudflare | Applications built on Cloudflare AI Search cannot retrieve your pages. If you are the one … |
| [DuckAssistBot](https://www.pathwren.workers.dev/crawler/duckassistbot.html) | `DuckAssistBot` | DuckDuckGo | No DuckAssist answers or citations from your site. Ordinary DuckDuckGo results are unaffec… |
| [ExaSearchBot](https://www.pathwren.workers.dev/crawler/exasearchbot.html) | `ExaSearchBot` | Exa | Agents built on Exa's API stop finding you. Exa publishes no statement about robots.txt co… |
| [Google-CloudVertexBot](https://www.pathwren.workers.dev/crawler/google-cloudvertexbot.html) | `Google-CloudVertexBot` | Google | Third parties can no longer build Vertex AI agents that read your site. Irrelevant to Goog… |
| [KlaviyoAIBot](https://www.pathwren.workers.dev/crawler/klaviyoaibot.html) | `KlaviyoAIBot` | Klaviyo | If the connected domain is yours, blocking this breaks the agent you configured. If it is … |
| [Meta-WebIndexer](https://www.pathwren.workers.dev/crawler/meta-webindexer.html) | `Meta-WebIndexer` | Meta | You leave the index Meta AI answers from across Facebook, Instagram and WhatsApp — the lar… |
| [OAI-SearchBot](https://www.pathwren.workers.dev/crawler/oai-searchbot.html) | `OAI-SearchBot` | OpenAI | High. Blocking this removes you from ChatGPT search results and from the source links Chat… |
| [PerplexityBot](https://www.pathwren.workers.dev/crawler/perplexitybot.html) | `PerplexityBot` | Perplexity | You stop being indexed and cited by Perplexity, and lose the referral clicks its citations… |
| [PhindBot](https://www.pathwren.workers.dev/crawler/phindbot.html) | `PhindBot` | Phind | You stop being cited in answers to technical questions — which, for documentation and refe… |
| [QualifiedBot](https://www.pathwren.workers.dev/crawler/qualifiedbot.html) | `QualifiedBot` | Qualified | A chatbot on a site that licensed the product loses context. If that site is yours, this b… |
| [ShapBot](https://www.pathwren.workers.dev/crawler/shapbot.html) | `ShapBot` | Parallel | Agents using Parallel's research API lose you as a source. Parallel documents robots.txt c… |
| [TerraCotta](https://www.pathwren.workers.dev/crawler/terracotta.html) | `TerraCotta` | Ceramic AI | You are absent from another agent-facing retrieval index. Ceramic documents that it obeys … |
| [YouBot](https://www.pathwren.workers.dev/crawler/youbot.html) | `YouBot` | You.com | Removal from You.com's index and from answers built on its API.… |

[json](https://www.pathwren.workers.dev/category/ai-search.json)

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/category/ai-search.html)
- [JSON](https://www.pathwren.workers.dev/category/ai-search.json)
- [Markdown](https://www.pathwren.workers.dev/category/ai-search.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/category/ai-search.html](https://www.pathwren.workers.dev/category/ai-search.html), generated from that page's own bytes in the same build. The HTML page is canonical.
