---
title: "MCP server — Agent Discovery Doctor"
description: "Check which of the 22 discovery documents agents actually ask for — llms.txt, A2A agent card, owners.json, oauth metadata, mcp.json, apis.json — a host serves, and who asks for each missing one. Streamable HTTP at /mcp/doctor, no key, no signup."
canonical: "https://www.pathwren.workers.dev/mcp-doctor.html"
url: "https://www.pathwren.workers.dev/mcp-doctor.md"
format: "markdown"
source: "the bytes of /mcp-doctor.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-03T07:55:35+00:00"
license: "CC0-1.0"
---

# Agent Discovery Doctor — MCP server

> Check which of the 22 discovery documents agents actually ask for — llms.txt, A2A agent card, owners.json, oauth metadata, mcp.json, apis.json — a host serves, and who asks for each missing one. Streamable HTTP at /mcp/doctor, no key, no signup.

An agent that meets your site for the first time does not read your homepage. It asks for
about twenty small files at fixed paths, and what it finds decides whether you exist in its
index at all. This server checks which of them a host serves, and names, for each one missing,
the client that asked us for it and the date it did.

```text
# which discovery documents does a host serve?
curl -s https://www.pathwren.workers.dev/mcp/doctor \
  -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"check_discovery_documents",
       "arguments":{"host":"example.com"}}}' \
  | jq -r '.result.structuredContent.documents[] | "\(.verdict)\t\(.path)"'

served      /robots.txt
missing     /llms.txt
missing     /.well-known/agent-card.json
soft-404    /.well-known/mcp.json
```

Each missing line comes back with who asks for it, when they asked here, and what the 404
costs — not a style-guide opinion:

```bash
curl -s https://www.pathwren.workers.dev/mcp/doctor \
  -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"explain_document",
       "arguments":{"name":"owners.json"}}}' | jq -r '.result.structuredContent.observed_askers[]
       | "\(.at)\t\(.ua)"'

2026-08-31T23:12:20Z	VerifyMCP-OwnersBot/1.0 (+https://verifymcp.io/docs/build/owners-json)
```

## Add it to a client

```text
claude mcp add --transport http agent-discovery-doctor https://www.pathwren.workers.dev/mcp/doctor
```

Clients that take a JSON config (Claude Desktop, Cursor, VS Code, Windsurf):

```json
{
 "mcpServers": {
  "agent-discovery-doctor": { "type": "streamable-http", "url": "https://www.pathwren.workers.dev/mcp/doctor" }
 }
}
```

## Tools

| Tool | What it answers |
| --- | --- |
| `check_discovery_documents` | The whole job. GETs the 22 paths on a host you name and returns each as served, missing, gated or soft-404 — a 200 carrying an HTML error page, which is worse than a 404 because the reader believes it — with who asks for each missing one. |
| `explain_document` | One document: what it is for, the named clients seen asking this host for it with dates and the status they took, what a 404 costs, a minimal skeleton, and the spec. No argument returns the whole catalogue. |
| `validate_llms_txt` | Paste an llms.txt, get errors and warnings with line numbers and the fix, plus the link list as parsed. Checks the format, not your prose. |
| `llms_txt_from_sitemap` | Paste sitemap.xml or a list of URLs, get a draft llms.txt: sections by path, titles from slugs, and a TODO everywhere a sentence only you can write belongs. |
| `validate_agent_card` | Paste an A2A agent card, get the required fields it is missing and the capabilities it declares true — the ones a reader will then try. |

## The 22 documents, and who actually asked

The catalogue is not a reading of the specs. Every row below is a request that arrived at
*this* host, with the user-agent as it came and the status it took:

| Document | Asked for here by | When | It got |
| --- | --- | --- | --- |
| `/.well-known/agent-card.json` | GolemreachTrustBot/0.1 | 2026-09-01 00:48Z | 404 — then 200 on its return at 01:59Z, once we shipped one |
| `/.well-known/agent.json` | GolemreachTrustBot/0.1 | 2026-09-01 00:48Z | 404 — it asks for both paths in the same second |
| `/.well-known/owners.json` | VerifyMCP-OwnersBot/1.0 | 2026-08-31 23:12Z | 404 at `/` and at `/mcp/`, in the same second |
| `/.well-known/oauth-protected-resource` | mcpbeat/0.1, exaforce-mcprep/0.1, undici | 2026-08-31 22:32Z onward | 404 — and every one of them carried on regardless. Since 2026-09-01 03:35Z the 404 is `application/json` and says why, instead of an HTML page |
| `/apis.json` and 11 more | apis.io-submit/1.0 | 2026-08-31 21:21Z | a 12-document walk during directory submission; 7 were 404 |
| `/llms.txt` | ClaudeBot/1.0 | 2026-09-01 01:10Z | 200 |
| `/.well-known/agent-card.json` | SaSameAgentAudit/0.1 | 2026-09-01 01:06Z | 404 |
| `/.well-known/x402` | AgenstryBot/0.3.0 | 2026-09-01 04:28Z | 404 — now 200. Payment discovery: `accepts` is empty because nothing here is paid, and the body says `implemented: false` so serving it is not mistaken for running the protocol |

The other documents in the catalogue — `ai.txt`, `api-catalog`,
`ai-plugin.json`, `swagger.json`, `security.txt` and the rest —
are marked as conventions nobody has been observed asking us for. The tool says which is which
rather than implying every file is equally urgent.

## What it will not do

**It refuses to check this host.** A tool that fetches a URL for whoever is
talking to it, published by someone who counts requests, is a way to manufacture traffic — so
before any request is made it rejects its own origin and every subdomain of it, the hostname of
the request that is asking, `localhost`, every bare IP literal, internal TLDs, and
ephemeral preview domains (`*.trycloudflare.com`, `*.ngrok.io`,
`*.vercel.app` previews). The refusal names the host and the reason. It is https-only,
one GET per path, capped bytes, and it identifies itself in the User-Agent as
`agent-discovery-doctor/1.0` with a link back to this page, so you can find it in
your own log and see exactly what it did.

The other four tools fetch nothing at all: text in, verdict out.

## How is this different from the other two?

[ai-crawler-index](https://www.pathwren.workers.dev/mcp.html) answers questions about crawlers.
[crawler-log-triage](https://www.pathwren.workers.dev/mcp-triage.html) reads a log you already have. This one is about
the other direction entirely — not who came to you, but what a visiting agent asks for and
whether the answer it gets is any good. No tool name, and no argument, is shared with either.

Protocol versions 2025-06-18, negotiated per call.
`server/discover` answers for clients on 2026-07-28, `initialize` for everyone else.
Read-only, stateless, no key. Listed in the
[official MCP Registry](https://registry.modelcontextprotocol.io/v0/servers?search=agent-discovery-doctor) as `dev.workers.pathwren.www/agent-discovery-doctor`.

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/mcp-doctor.html)
- [JSON](https://www.pathwren.workers.dev/mcp-doctor.json)
- [Markdown](https://www.pathwren.workers.dev/mcp-doctor.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/mcp-doctor.html](https://www.pathwren.workers.dev/mcp-doctor.html), generated from that page's own bytes in the same build. The HTML page is canonical.
