---
title: "AI training crawlers (27) — AI Crawler Index"
description: "Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today."
canonical: "https://www.pathwren.workers.dev/category/ai-training.html"
url: "https://www.pathwren.workers.dev/category/ai-training.md"
format: "markdown"
source: "the bytes of /category/ai-training.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T06:29:54+00:00"
license: "CC0-1.0"
---

# AI training crawlers

> Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today.

Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today.

| Crawler | Token | Operator | Cost of blocking |
| --- | --- | --- | --- |
| [anthropic-ai](https://www.pathwren.workers.dev/crawler/anthropic-ai.html) | `anthropic-ai` | Anthropic | None. Nothing crawls under this name today; keeping the rule is harmless insurance.… |
| [Applebot-Extended](https://www.pathwren.workers.dev/crawler/applebot-extended.html) | `Applebot-Extended` | Apple | Excluded from Apple Intelligence training. Siri, Spotlight and Safari suggestions are unaf… |
| [Bytespider](https://www.pathwren.workers.dev/crawler/bytespider.html) | `Bytespider` | ByteDance | Little to lose. If you want it gone, expect to block by user-agent at the edge rather than… |
| [ClaudeBot](https://www.pathwren.workers.dev/crawler/claudebot.html) | `ClaudeBot` | Anthropic | Content excluded from training data for future Claude models. No effect on Claude's abilit… |
| [cohere-training-data-crawler](https://www.pathwren.workers.dev/crawler/cohere-training-data-crawler.html) | `cohere-training-data-crawler` | Cohere | Excluded from Cohere model training.… |
| [Cotoyogi](https://www.pathwren.workers.dev/crawler/cotoyogi.html) | `Cotoyogi` | ROIS-DS | Your Japanese-language content is left out of an academic training corpus.… |
| [FacebookBot](https://www.pathwren.workers.dev/crawler/facebookbot.html) | `FacebookBot` | Meta | Negligible today. Keep the rule; expect little traffic.… |
| [Factset_spyderbot](https://www.pathwren.workers.dev/crawler/factset-spyderbot.html) | `Factset_spyderbot` | FactSet | Exclusion from a financial-data vendor's corpus. Relevant mostly to companies whose filing… |
| [Google-Extended](https://www.pathwren.workers.dev/crawler/google-extended.html) | `Google-Extended` | Google | You are excluded from Gemini grounding and Gemini training. Google Search ranking and inde… |
| [GoogleOther](https://www.pathwren.workers.dev/crawler/googleother.html) | `GoogleOther` | Google | No effect on Search indexing. Blocks internal Google research and product fetches.… |
| [GoogleOther-Image](https://www.pathwren.workers.dev/crawler/googleother-image.html) | `GoogleOther-Image` | Google | Google teams outside Search stop fetching your images. Image Search itself is unaffected —… |
| [GoogleOther-Video](https://www.pathwren.workers.dev/crawler/googleother-video.html) | `GoogleOther-Video` | Google | No effect on Search or on Google Video search. Blocks internal Google research fetches of … |
| [GPTBot](https://www.pathwren.workers.dev/crawler/gptbot.html) | `GPTBot` | OpenAI | Your content is excluded from training data for future OpenAI models. No effect on ChatGPT… |
| [ICC-Crawler](https://www.pathwren.workers.dev/crawler/icc-crawler.html) | `ICC-Crawler` | NICT | You are excluded from a national research corpus and from the commercial redistributions o… |
| [ISSCyberRiskCrawler](https://www.pathwren.workers.dev/crawler/isscyberriskcrawler.html) | `ISSCyberRiskCrawler` | ISS Corporate Solutions | A rule here is a statement of intent. Your organisation's public footprint still gets scor… |
| [Linguee Bot](https://www.pathwren.workers.dev/crawler/linguee-bot.html) | `Linguee Bot` | Linguee | Multilingual pages stop feeding a translation corpus. If your site is translated, being in… |
| [meta-externalagent](https://www.pathwren.workers.dev/crawler/meta-externalagent.html) | `meta-externalagent` | Meta | Excluded from Meta AI training. Link previews on Facebook, Instagram and WhatsApp are unaf… |
| [Poseidon Research Crawler](https://www.pathwren.workers.dev/crawler/poseidon-research-crawler.html) | `Poseidon Research Crawler` | Poseidon Research | Exclusion from an interpretability research corpus. No published compliance statement.… |
| [QuillBot](https://www.pathwren.workers.dev/crawler/quillbot.html) | `QuillBot` | QuillBot | Exclusion from QuillBot's corpus. No compliance statement is published, so the rule is a r… |
| [Reflectionbot](https://www.pathwren.workers.dev/crawler/reflectionbot.html) | `Reflectionbot` | Reflection AI | Unknown by construction — which is itself the reason some people block it. Nothing user-fa… |
| [SBIntuitionsBot](https://www.pathwren.workers.dev/crawler/sbintuitionsbot.html) | `SBIntuitionsBot` | SB Intuitions | Your content is excluded from a Japanese-language foundation-model corpus. Nothing user-fa… |
| [SemrushBot-OCOB](https://www.pathwren.workers.dev/crawler/semrushbot-ocob.html) | `SemrushBot-OCOB` | Semrush | Exclusion from Semrush's AI corpus, with its SEO crawl unaffected.… |
| [Sidetrade indexer bot](https://www.pathwren.workers.dev/crawler/sidetrade-indexer-bot.html) | `Sidetrade indexer bot` | Sidetrade | Exclusion from a commercial B2B dataset. The operator publishes no robots.txt statement, s… |
| [TikTokSpider](https://www.pathwren.workers.dev/crawler/tiktokspider.html) | `TikTokSpider` | ByteDance | Little to lose unless TikTok search referral matters to you.… |
| [Webzio-Extended](https://www.pathwren.workers.dev/crawler/webzio-extended.html) | `Webzio-Extended` | Webz.io | Your content is excluded from the AI-training tier of Webz.io's product while ordinary col… |
| [YandexAdditional](https://www.pathwren.workers.dev/crawler/yandexadditional.html) | `YandexAdditional` | Yandex | You disappear from Yandex's AI answers while staying in Yandex Search. This is Yandex's eq… |
| [YandexAdditionalBot](https://www.pathwren.workers.dev/crawler/yandexadditionalbot.html) | `YandexAdditionalBot` | Yandex | Same as YandexAdditional: out of Yandex's AI answers, still in Yandex Search. Name both to… |

[json](https://www.pathwren.workers.dev/category/ai-training.json)

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/category/ai-training.html)
- [JSON](https://www.pathwren.workers.dev/category/ai-training.json)
- [Markdown](https://www.pathwren.workers.dev/category/ai-training.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/category/ai-training.html](https://www.pathwren.workers.dev/category/ai-training.html), generated from that page's own bytes in the same build. The HTML page is canonical.
