---
title: "robots.txt: Block the crawlers with disputed robots compliance — AI Crawler Index"
description: "The ones repeatedly reported as ignoring robots.txt. Included for completeness — expect to enforce this at the edge instead."
canonical: "https://www.pathwren.workers.dev/policy/block-disputed.html"
url: "https://www.pathwren.workers.dev/policy/block-disputed.md"
format: "markdown"
source: "the bytes of /policy/block-disputed.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T06:29:54+00:00"
license: "CC0-1.0"
---

# Block the crawlers with disputed robots compliance

> The ones repeatedly reported as ignoring robots.txt. Included for completeness — expect to enforce this at the edge instead.

The ones repeatedly reported as ignoring robots.txt. Included for completeness — expect to enforce this at the edge instead.

```bash
curl -s https://www.pathwren.workers.dev/robots/block-disputed.txt >> robots.txt
```

A robots.txt rule is a request. For the operators in this file the request is documented as unreliable or explicitly not applicable, so the honest use of this file is as a record of intent that sits alongside a real block by user-agent or by IP at your CDN.

## Names 18 crawlers

[Bytespider](https://www.pathwren.workers.dev/crawler/bytespider.html) · [FeedFetcher-Google](https://www.pathwren.workers.dev/crawler/feedfetcher-google.html) · [Google-Agent](https://www.pathwren.workers.dev/crawler/google-agent.html) · [Google-CWS](https://www.pathwren.workers.dev/crawler/google-cws.html) · [Google-GeminiNotebook](https://www.pathwren.workers.dev/crawler/google-gemininotebook.html) · [Google-Pinpoint](https://www.pathwren.workers.dev/crawler/google-pinpoint.html) · [Google-Read-Aloud](https://www.pathwren.workers.dev/crawler/google-read-aloud.html) · [Google-Safety](https://www.pathwren.workers.dev/crawler/google-safety.html) · [Google-Site-Verification](https://www.pathwren.workers.dev/crawler/google-site-verification.html) · [GoogleMessages](https://www.pathwren.workers.dev/crawler/googlemessages.html) · [GoogleProducer](https://www.pathwren.workers.dev/crawler/googleproducer.html) · [ISSCyberRiskCrawler](https://www.pathwren.workers.dev/crawler/isscyberriskcrawler.html) · [LAIONDownloader](https://www.pathwren.workers.dev/crawler/laiondownloader.html) · [Linguee Bot](https://www.pathwren.workers.dev/crawler/linguee-bot.html) · [Perplexity-User](https://www.pathwren.workers.dev/crawler/perplexity-user.html) · [Thinkbot](https://www.pathwren.workers.dev/crawler/thinkbot.html) · [TikTokSpider](https://www.pathwren.workers.dev/crawler/tiktokspider.html) · [YandexMetrika](https://www.pathwren.workers.dev/crawler/yandexmetrika.html)

## The file

[/robots/block-disputed.txt](https://www.pathwren.workers.dev/robots/block-disputed.txt) · [json](https://www.pathwren.workers.dev/policy/block-disputed.json)

```text
# AI Crawler Index — policy: block-disputed
# Block the crawlers with disputed robots compliance
# The ones repeatedly reported as ignoring robots.txt. Included for completeness — expect to enforce this at the edge instead.
# Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/block-disputed.html
# 18 crawlers named. Paste into robots.txt at your document root.

User-agent: Bytespider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: FeedFetcher-Google   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Agent   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-CWS   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-GeminiNotebook   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Pinpoint   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Read-Aloud   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Safety   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Google-Site-Verification   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: GoogleMessages   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: GoogleProducer   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: ISSCyberRiskCrawler   # compliance disputed; enforce at the edge
Disallow: /

User-agent: LAIONDownloader   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Linguee Bot   # compliance disputed; enforce at the edge
Disallow: /

User-agent: Perplexity-User   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: Thinkbot   # compliance disputed; enforce at the edge
Disallow: /

User-agent: TikTokSpider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: YandexMetrika   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml
```

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/policy/block-disputed.html)
- [JSON](https://www.pathwren.workers.dev/policy/block-disputed.json)
- [Markdown](https://www.pathwren.workers.dev/policy/block-disputed.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/policy/block-disputed.html](https://www.pathwren.workers.dev/policy/block-disputed.html), generated from that page's own bytes in the same build. The HTML page is canonical.
