curl -s https://www.pathwren.workers.dev/crawler/ccbot.json # this page, as JSON
No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a
Common Crawl · Corpus and dataset builders · json
User-agent: CCBot Disallow: /
| robots.txt token | CCBot |
| User-agent contains | CCBot |
| Operator | Common Crawl |
| Category | Corpus and dataset builders |
| robots.txt | obeys robots.txt (documented) |
| Verify by | published IP ranges |
| Published ranges | 4 IPv4 + 1 IPv6 · json · source |
Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list.
Future Common Crawl snapshots exclude you, so downstream training sets lose you too — but only going forward. Existing snapshots are permanent and blocking today does not retract them.
CCBot/2.0 (https://commoncrawl.org/faq/)
User-agent: CCBot Allow: /
Operator documentation: https://commoncrawl.org/ccbot
Machine copies: json ·
markdown
Policies that name this crawler:
allow-all · block-all-ai · block-datasets · maximum-ai-visibility