---
title: "robots.txt: Block corpus and dataset builders — AI Crawler Index"
description: "Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot."
canonical: "https://www.pathwren.workers.dev/policy/block-datasets.html"
url: "https://www.pathwren.workers.dev/policy/block-datasets.md"
format: "markdown"
source: "the bytes of /policy/block-datasets.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T06:29:54+00:00"
license: "CC0-1.0"
---

# Block corpus and dataset builders

> Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot.

Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot.

```bash
curl -s https://www.pathwren.workers.dev/robots/block-datasets.txt >> robots.txt
```

These are the highest-leverage blocks per line, because one crawl becomes many downstream training runs. It is also the block with the longest delay before it has any effect, and no effect at all on archives already published.

## Names 17 crawlers

[AI2Bot](https://www.pathwren.workers.dev/crawler/ai2bot.html) · [Ai2Bot-Dolma](https://www.pathwren.workers.dev/crawler/ai2bot-dolma.html) · [aiHitBot](https://www.pathwren.workers.dev/crawler/aihitbot.html) · [AwarioRssBot](https://www.pathwren.workers.dev/crawler/awariorssbot.html) · [AwarioSmartBot](https://www.pathwren.workers.dev/crawler/awariosmartbot.html) · [CCBot](https://www.pathwren.workers.dev/crawler/ccbot.html) · [Diffbot](https://www.pathwren.workers.dev/crawler/diffbot.html) · [EchoboxBot](https://www.pathwren.workers.dev/crawler/echoboxbot.html) · [ImagesiftBot](https://www.pathwren.workers.dev/crawler/imagesiftbot.html) · [img2dataset](https://www.pathwren.workers.dev/crawler/img2dataset.html) · [LAIONDownloader](https://www.pathwren.workers.dev/crawler/laiondownloader.html) · [omgili](https://www.pathwren.workers.dev/crawler/omgili.html) · [omgilibot](https://www.pathwren.workers.dev/crawler/omgilibot.html) · [Panscient](https://www.pathwren.workers.dev/crawler/panscient.html) · [Thinkbot](https://www.pathwren.workers.dev/crawler/thinkbot.html) · [VelenPublicWebCrawler](https://www.pathwren.workers.dev/crawler/velenpublicwebcrawler.html) · [YaK](https://www.pathwren.workers.dev/crawler/yak.html)

## The file

[/robots/block-datasets.txt](https://www.pathwren.workers.dev/robots/block-datasets.txt) · [json](https://www.pathwren.workers.dev/policy/block-datasets.json)

```text
# AI Crawler Index — policy: block-datasets
# Block corpus and dataset builders
# Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot.
# Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/block-datasets.html
# 17 crawlers named. Paste into robots.txt at your document root.

User-agent: AI2Bot
Disallow: /

User-agent: Ai2Bot-Dolma
Disallow: /

User-agent: aiHitBot
Disallow: /

User-agent: AwarioRssBot
Disallow: /

User-agent: AwarioSmartBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: EchoboxBot
Disallow: /

User-agent: ImagesiftBot
Disallow: /

User-agent: img2dataset
Disallow: /

User-agent: LAIONDownloader   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: omgili
Disallow: /

User-agent: omgilibot
Disallow: /

User-agent: panscient.com
Disallow: /

User-agent: Thinkbot   # compliance disputed; enforce at the edge
Disallow: /

User-agent: VelenPublicWebCrawler
Disallow: /

User-agent: YaK
Disallow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml
```

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/policy/block-datasets.html)
- [JSON](https://www.pathwren.workers.dev/policy/block-datasets.json)
- [Markdown](https://www.pathwren.workers.dev/policy/block-datasets.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/policy/block-datasets.html](https://www.pathwren.workers.dev/policy/block-datasets.html), generated from that page's own bytes in the same build. The HTML page is canonical.
