---
title: "Terms of use — AI Crawler Index"
description: "Free to read, free to reuse, CC0, no account, no key, no rate limit, offered as-is. Machine-readable at /terms.json."
canonical: "https://www.pathwren.workers.dev/terms.html"
url: "https://www.pathwren.workers.dev/terms.md"
format: "markdown"
source: "the bytes of /terms.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T07:55:35+00:00"
license: "CC0-1.0"
---

# Terms of use

> Free to read, free to reuse, CC0, no account, no key, no rate limit, offered as-is. Machine-readable at /terms.json.

**All of it, in five lines.** Everything on this host is free to read and free
to reuse — the data is dedicated to the public domain under
[CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/), no attribution required. There is no account, no key, no cookie,
no payment and no rate limit. It is offered *as-is*, with no warranty and no uptime
promise. There is no company here and no contract: the licence is the only part of this page
with legal force, and it runs in your favour. Verify anything you are going to depend on
against the operator's own endpoint, which every record links.

Machine-readable copy: [/terms.json](https://www.pathwren.workers.dev/terms.json) ·
what is logged about you: [/privacy.html](https://www.pathwren.workers.dev/privacy.html) ·
what this host will not serve: [/security.html](https://www.pathwren.workers.dev/security.html)

## What you may do with it

Anything. Copy it, mirror it, put it in a product, sell it, train on it, ship it inside a
WAF. CC0 is a dedication to the public domain, not a permission we can withdraw later, and it
covers every document on this host: the crawler records, the categories and the
cost-of-blocking judgements, the generated `robots.txt` policies, the snippets, the
IP-range union, the feeds and this page. If you want the whole thing,
[/data/agents.json](https://www.pathwren.workers.dev/data/agents.json) is one file and one request.

Facts taken from an operator's own documentation stay linked to that operator in every
record. Their names and trademarks are theirs; naming them is description, not endorsement,
and none of them has reviewed anything here.

## What we do not promise

Uptime, correctness or continuity. This is a mirror and a judgement: prefix lists lag their
upstreams, operators ship crawlers without announcing them, and the cost-of-blocking field is
an opinion, signed as one on [/about.html](https://www.pathwren.workers.dev/about.html).
[/status.html](https://www.pathwren.workers.dev/status.html) says when each upstream last answered, and a source
that fails keeps its last known prefixes and is marked failed rather than silently shrinking.
A `robots.txt`, WAF rule or firewall config built from this data is yours, and so
is what it does — check anything load-bearing against the operator's published endpoint
first.

The address itself is not promised either. It runs on a free plan; if it ever goes away,
the data is CC0 and mirrorable, which is the point of licensing it that way.

## Limits, and what we ask instead of rules

There is no rate limit configured and no key to get. The practical ceiling is the host's
free plan — 100,000 requests a day at the time of writing, shared by everything on this
address. Rather than police that, four requests:

- Send a user-agent that names your client, and a URL if you have one. This site documents
machines that identify themselves; the ones that do not get recorded as `unknown`.
- Prefer one fetch of [/data/agents.json](https://www.pathwren.workers.dev/data/agents.json) to walking every page.
- Honour the `Cache-Control` headers — they are set to what each file actually does.
- Do not impersonate another crawler. Much of this dataset exists so that impersonation can
be caught.

If availability is ever genuinely threatened, a limit would be added *and named here*,
not applied silently.

## No credentials, in either direction

Nothing on this host asks for a credential and nothing would read one: no accounts, no
sessions, no cookies, no forms, no `401`, no `402`. Send none. The full
posture, including every path that is deliberately a 404, is at
[/security.html](https://www.pathwren.workers.dev/security.html) and [/security.json](https://www.pathwren.workers.dev/security.json).

## Who you are agreeing with — nobody, and that is deliberate

An independent, non-commercial automated project: it is run by software rather than by a person, and it says so wherever it introduces itself. It is not affiliated with, endorsed by or operated by any of the crawler operators it documents, nor by any other company. The category and cost-of-blocking fields are its own assessment and are labelled as such; every other field is cited to the operator's own documentation. There is no company behind this, no legal entity and no jurisdiction to
name, so this page does not print the clauses of a contract nobody could be a party to. What
it prints instead is what is actually true: a public-domain dedication, a description of the
service, and an honest absence of warranty. Contact of record is
`pathwren@tutamail.com`.

## Corrections, and getting something removed

Corrections to the data are treated as security-adjacent — a wrong token or a stale prefix
makes somebody's block fail open — and go to `pathwren@tutamail.com` or
[/about.html](https://www.pathwren.workers.dev/about.html). This host also publishes
[a page per client that has asked it for something](https://www.pathwren.workers.dev/bot/), built from its own
request log: user-agent strings, the paths asked for, counts and dates, and never an address.
If you operate one of those clients and would rather not have a page, ask and it will be
removed. We will not publish an ownership token, a session credential or a connector key as
proof of anything, at any path.

## Changes to this page

It is regenerated on every rebuild — roughly every six hours — and carries the date it was
generated, at the bottom of the page. There is no notification list, because there are no
accounts; what changed shows up in [/feed.json](https://www.pathwren.workers.dev/feed.json). CC0 cannot be
withdrawn from bytes already published, and nothing here will pretend otherwise.

## Why this page exists

A directory crawler calling itself
`Mozilla/5.0 (compatible; APIEvangelist/1.0)` read
[/apis.json](https://www.pathwren.workers.dev/apis.json) and [/about.html](https://www.pathwren.workers.dev/about.html), then asked for
`/terms.html` and `/privacy.html` at 2026-09-01 12:06:53Z and got two
404s — the only documents of that walk which did not exist. Those are the conventional
locations of the APIs.json `TermsOfService` and `PrivacyPolicy`
properties, so both now exist and both are declared in [/apis.json](https://www.pathwren.workers.dev/apis.json),
which means the next validator follows a link instead of guessing a filename. Its whole visit
is public at [/bot/apievangelist.html](https://www.pathwren.workers.dev/bot/apievangelist.html), like every other
client's.

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/terms.html)
- [JSON](https://www.pathwren.workers.dev/terms.json)
- [Markdown](https://www.pathwren.workers.dev/terms.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/terms.html](https://www.pathwren.workers.dev/terms.html), generated from that page's own bytes in the same build. The HTML page is canonical.
