curl -s https://www.pathwren.workers.dev/crawler/ccbot.json   # this page, as JSON

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

CCBot

Common Crawl · Corpus and dataset builders · json

User-agent: CCBot
Disallow: /
robots.txt tokenCCBot
User-agent containsCCBot
OperatorCommon Crawl
CategoryCorpus and dataset builders
robots.txtobeys robots.txt (documented)
Verify bypublished IP ranges
Published ranges4 IPv4 + 1 IPv6 · json · source

What it is

Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list.

What blocking it costs you

Future Common Crawl snapshots exclude you, so downstream training sets lose you too — but only going forward. Existing snapshots are permanent and blocking today does not retract them.

Full user-agent string

CCBot/2.0 (https://commoncrawl.org/faq/)

Allow it instead

User-agent: CCBot
Allow: /

Operator documentation: https://commoncrawl.org/ccbot
Machine copies: json · markdown
Policies that name this crawler: allow-all · block-all-ai · block-datasets · maximum-ai-visibility