# AI Crawler Index — policy: block-datasets # Block corpus and dataset builders # Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot. # Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/block-datasets.html # 17 crawlers named. Paste into robots.txt at your document root. User-agent: AI2Bot Disallow: / User-agent: Ai2Bot-Dolma Disallow: / User-agent: aiHitBot Disallow: / User-agent: AwarioRssBot Disallow: / User-agent: AwarioSmartBot Disallow: / User-agent: CCBot Disallow: / User-agent: Diffbot Disallow: / User-agent: EchoboxBot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: img2dataset Disallow: / User-agent: LAIONDownloader # operator states robots.txt does not apply; enforce at the edge Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: panscient.com Disallow: / User-agent: Thinkbot # compliance disputed; enforce at the edge Disallow: / User-agent: VelenPublicWebCrawler Disallow: / User-agent: YaK Disallow: / User-agent: * Allow: / Sitemap: https://www.pathwren.workers.dev/sitemap.xml