curl -s https://www.pathwren.workers.dev/crawler/googlebot-news.json # this page, as JSON
No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a
Google · Search engines · json
User-agent: Googlebot-News Disallow: /
| robots.txt token | Googlebot-News |
| User-agent contains | Googlebot-News |
| Operator | |
| Category | Search engines |
| robots.txt | obeys robots.txt (documented) |
| Verify by | published IP ranges |
| Published ranges | 170 IPv4 + 147 IPv6 · json · source |
A robots.txt token controlling inclusion in Google News. It does not have its own user-agent string; the fetch arrives as Googlebot.
Removal from Google News, with normal Search unaffected.
(uses the Googlebot user-agent; controlled by the Googlebot-News robots token)
User-agent: Googlebot-News Allow: /
Operator documentation: https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers
Machine copies: json ·
markdown
Policies that name this crawler:
allow-all · allow-ai-search-only · maximum-ai-visibility