curl -s https://www.pathwren.workers.dev/crawler/gptbot.json # this page, as JSON
No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a
OpenAI · AI training crawlers · json
User-agent: GPTBot Disallow: /
| robots.txt token | GPTBot |
| User-agent contains | GPTBot |
| Operator | OpenAI |
| Category | AI training crawlers |
| robots.txt | obeys robots.txt (documented) |
| Verify by | published IP ranges |
| Published ranges | 21 IPv4 + 0 IPv6 · json · source |
OpenAI's bulk crawler. Pages it fetches may be used to train future OpenAI foundation models. It is not the bot that puts you in ChatGPT's search results, and blocking it does not remove you from them.
Your content is excluded from training data for future OpenAI models. No effect on ChatGPT search visibility, on citations, or on links a user pastes into ChatGPT.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot
User-agent: GPTBot Allow: /
Operator documentation: https://platform.openai.com/docs/bots
Machine copies: json ·
markdown
Policies that name this crawler:
allow-all · block-ai-training · block-all-ai · maximum-ai-visibility