curl -s https://www.pathwren.workers.dev/crawler/cohere-training-data-crawler.json   # this page, as JSON

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

cohere-training-data-crawler

Cohere · AI training crawlers · json

User-agent: cohere-training-data-crawler
Disallow: /
robots.txt tokencohere-training-data-crawler
User-agent containscohere-training-data-crawler
OperatorCohere
CategoryAI training crawlers
robots.txtobeys robots.txt (documented)
Verify byno published verification method

What it is

Cohere's separately-named bulk crawler for model training data, split out so consent for training and consent for retrieval can differ.

What blocking it costs you

Excluded from Cohere model training.

Full user-agent string

Mozilla/5.0 (compatible; cohere-training-data-crawler)

Allow it instead

User-agent: cohere-training-data-crawler
Allow: /

Operator documentation: https://cohere.com/
Machine copies: json · markdown
Policies that name this crawler: allow-all · block-ai-training · block-all-ai · maximum-ai-visibility