---
title: "Corpus and dataset builders (17) — AI Crawler Index"
description: "Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect."
canonical: "https://www.pathwren.workers.dev/category/dataset.html"
url: "https://www.pathwren.workers.dev/category/dataset.md"
format: "markdown"
source: "the bytes of /category/dataset.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T06:29:54+00:00"
license: "CC0-1.0"
---

# Corpus and dataset builders

> Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect.

Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect.

| Crawler | Token | Operator | Cost of blocking |
| --- | --- | --- | --- |
| [AI2Bot](https://www.pathwren.workers.dev/crawler/ai2bot.html) | `AI2Bot` | Allen Institute for AI | Excluded from open research datasets. Worth a deliberate decision: this is the category wh… |
| [Ai2Bot-Dolma](https://www.pathwren.workers.dev/crawler/ai2bot-dolma.html) | `Ai2Bot-Dolma` | Allen Institute for AI | Same as AI2Bot: exclusion from an open, published training corpus.… |
| [aiHitBot](https://www.pathwren.workers.dev/crawler/aihitbot.html) | `aiHitBot` | aiHit | Your company record in a B2B dataset goes stale. Documented as respecting robots.txt, so t… |
| [AwarioRssBot](https://www.pathwren.workers.dev/crawler/awariorssbot.html) | `AwarioRssBot` | Awario | Your RSS updates stop reaching Awario's monitoring. Block both tokens or neither.… |
| [AwarioSmartBot](https://www.pathwren.workers.dev/crawler/awariosmartbot.html) | `AwarioSmartBot` | Awario | Mentions of brands on your pages stop being surfaced to the people monitoring them — inclu… |
| [CCBot](https://www.pathwren.workers.dev/crawler/ccbot.html) | `CCBot` | Common Crawl | Future Common Crawl snapshots exclude you, so downstream training sets lose you too — but … |
| [Diffbot](https://www.pathwren.workers.dev/crawler/diffbot.html) | `Diffbot` | Diffbot | Your facts stop entering a widely-licensed knowledge graph. Whether that is a loss depends… |
| [EchoboxBot](https://www.pathwren.workers.dev/crawler/echoboxbot.html) | `EchoboxBot` | Echobox | Publishers using Echobox get worse scheduling decisions about your articles. No compliance… |
| [ImagesiftBot](https://www.pathwren.workers.dev/crawler/imagesiftbot.html) | `ImagesiftBot` | Hive AI | Your images stop entering an image dataset and reverse-image index.… |
| [img2dataset](https://www.pathwren.workers.dev/crawler/img2dataset.html) | `img2dataset` | LAION / img2dataset | Your images are skipped when someone materialises an image-text dataset that references th… |
| [LAIONDownloader](https://www.pathwren.workers.dev/crawler/laiondownloader.html) | `LAIONDownloader` | LAION / img2dataset | Your media is skipped when an open research dataset is built from URL lists. Once a datase… |
| [omgili](https://www.pathwren.workers.dev/crawler/omgili.html) | `omgili` | Webz.io | Same as omgilibot.… |
| [omgilibot](https://www.pathwren.workers.dev/crawler/omgilibot.html) | `omgilibot` | Webz.io | Exclusion from a commercial dataset resold to third parties.… |
| [Panscient](https://www.pathwren.workers.dev/crawler/panscient.html) | `panscient.com` | Panscient | Your company pages stop feeding a business-data product. No effect on search or assistants… |
| [Thinkbot](https://www.pathwren.workers.dev/crawler/thinkbot.html) | `Thinkbot` | Thinkbot | Exclusion from a market-research dataset. Expect to enforce this at the edge rather than i… |
| [VelenPublicWebCrawler](https://www.pathwren.workers.dev/crawler/velenpublicwebcrawler.html) | `VelenPublicWebCrawler` | Hunter (Velen) | Your company pages stop feeding a B2B contact and company dataset. The crawl rate it docum… |
| [YaK](https://www.pathwren.workers.dev/crawler/yak.html) | `YaK` | Meltwater | Your content stops appearing in Meltwater's media monitoring — which is how PR teams find … |

[json](https://www.pathwren.workers.dev/category/dataset.json)

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/category/dataset.html)
- [JSON](https://www.pathwren.workers.dev/category/dataset.json)
- [Markdown](https://www.pathwren.workers.dev/category/dataset.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/category/dataset.html](https://www.pathwren.workers.dev/category/dataset.html), generated from that page's own bytes in the same build. The HTML page is canonical.
