curl -s https://www.pathwren.workers.dev/crawler/icc-crawler.json # this page, as JSON
No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a
NICT · AI training crawlers · json
User-agent: ICC-Crawler Disallow: /
| robots.txt token | ICC-Crawler |
| User-agent contains | ICC-Crawler |
| Operator | NICT |
| Category | AI training crawlers |
| robots.txt | obeys robots.txt (documented) |
| Verify by | no published verification method |
Operated by NICT, Japan's national information and communications research institute. The collected data supports AI research and, per the operator, is also provided to third parties including commercial companies.
You are excluded from a national research corpus and from the commercial redistributions of it. This is a dataset-shaped block: one refusal, many downstream effects.
ICC-Crawler
User-agent: ICC-Crawler Allow: /
Operator documentation: https://www.nict.go.jp/en/
Machine copies: json ·
markdown
Policies that name this crawler:
allow-all · block-ai-training · block-all-ai · maximum-ai-visibility