curl -s https://www.pathwren.workers.dev/policy/allow-all.json   # this page, as JSON

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

Allow everything, explicitly

Every crawler on this index is named and allowed. Use when you want maximum reach into search and assistants and have nothing to withhold.

curl -s https://www.pathwren.workers.dev/robots/allow-all.txt >> robots.txt

An empty robots.txt already allows everything, so this file is not about permission — it is about being explicit. Naming each token means a later change is a one-line diff instead of a rewrite, and it documents that the allow was a decision. This is the policy this site itself serves.

Names 150 crawlers

AdsBot-Google · AdsBot-Google-Mobile · AdsBot-Google-Mobile-Apps · AhrefsBot · AhrefsSiteAudit · AI2Bot · Ai2Bot-Dolma · aiHitBot · AIWebIndex · Amazonbot · Andibot · Anomura · anthropic-ai · APIs-Google · Applebot · Applebot-Extended · archive.org_bot · atlassian-bot · AwarioRssBot · AwarioSmartBot · Baiduspider · Barkrowler · bedrockbot · bingbot · Bytespider · CCBot · ChatGPT Agent · ChatGPT-User · Claude-SearchBot · Claude-User · Claude-Web · ClaudeBot · Cloudflare-AutoRAG · cohere-ai · cohere-training-data-crawler · Cotoyogi · Crawl4AI · Crawlspace · DataForSeoBot · Diffbot · DotBot · DuckAssistBot · DuckDuckBot · EchoboxBot · ExaSearchBot · FacebookBot · facebookexternalhit · Factset_spyderbot · FeedFetcher-Google · FirecrawlAgent · Google-Agent · Google-CloudVertexBot · Google-CWS · Google-Extended · Google-GeminiNotebook · Google-InspectionTool · Google-Pinpoint · Google-Read-Aloud · Google-Safety · Google-Site-Verification · Googlebot · Googlebot-Image · Googlebot-News · Googlebot-Video · GoogleMessages · GoogleOther · GoogleOther-Image · GoogleOther-Video · GoogleProducer · GPTBot · ia_archiver · ICC-Crawler · ImagesiftBot · img2dataset · ISSCyberRiskCrawler · Kagibot · KlaviyoAIBot · LAIONDownloader · Lightpanda · Linguee Bot · Mediapartners-Google · meta-externalagent · meta-externalfetcher · Meta-WebIndexer · MistralAI-User · MJ12bot · MojeekBot · OAI-SearchBot · omgili · omgilibot · Panscient · Perplexity-User · PerplexityBot · PetalBot · PhindBot · Pinterestbot · Poseidon Research Crawler · QualifiedBot · QuillBot · Qwantbot · Qwantbot-news · Reflectionbot · rogerbot · SBIntuitionsBot · Scrapy · Screaming Frog SEO Spider · SemrushBot · SemrushBot-BA · SemrushBot-ESI · SemrushBot-FT · SemrushBot-OCOB · SemrushBot-SI · SemrushBot-SWA · SEOkicks · serpstatbot · SeznamBot · ShapBot · Sidetrade indexer bot · SiteAuditBot · Slackbot · Slackbot-LinkExpanding · SplitSignalBot · Storebot-Google · TerraCotta · Thinkbot · TikTokSpider · Timpibot · VelenPublicWebCrawler · Webzio-Extended · wpbot · YaK · YandexAdditional · YandexAdditionalBot · YandexBlogs · YandexBot · YandexCalendar · YandexComBot · YandexDirect · YandexFavicons · YandexImages · YandexMarket · YandexMedia · YandexMetrika · YandexMobileBot · YandexRenderResourcesBot · YandexScreenshotBot · YandexVideo · YandexWebmaster · Yeti · YouBot

The file

/robots/allow-all.txt · json

# AI Crawler Index — policy: allow-all
# Allow everything, explicitly
# Every crawler on this index is named and allowed. Use when you want maximum reach into search and assistants and have nothing to withhold.
# Generated 2026-09-03 from https://www.pathwren.workers.dev/policy/allow-all.html
# 150 crawlers named. Paste into robots.txt at your document root.

User-agent: AdsBot-Google
Allow: /

User-agent: AdsBot-Google-Mobile
Allow: /

User-agent: AdsBot-Google-Mobile-Apps
Allow: /

User-agent: AhrefsBot
Allow: /

User-agent: AhrefsSiteAudit
Allow: /

User-agent: AI2Bot
Allow: /

User-agent: Ai2Bot-Dolma
Allow: /

User-agent: aiHitBot
Allow: /

User-agent: AIWebIndex
Allow: /

User-agent: Amazonbot
Allow: /

User-agent: Andibot
Allow: /

User-agent: Anomura
Allow: /

User-agent: anthropic-ai   # control token, no crawler uses this user-agent
Allow: /

User-agent: APIs-Google
Allow: /

User-agent: Applebot
Allow: /

User-agent: Applebot-Extended   # control token, no crawler uses this user-agent
Allow: /

User-agent: archive.org_bot
Allow: /

User-agent: atlassian-bot
Allow: /

User-agent: AwarioRssBot
Allow: /

User-agent: AwarioSmartBot
Allow: /

User-agent: Baiduspider
Allow: /

User-agent: barkrowler
Allow: /

User-agent: bedrockbot
Allow: /

User-agent: bingbot
Allow: /

User-agent: Bytespider   # compliance disputed; enforce at the edge
Allow: /

User-agent: CCBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Claude-Web   # control token, no crawler uses this user-agent
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Cloudflare-AutoRAG
Allow: /

User-agent: cohere-ai
Allow: /

User-agent: cohere-training-data-crawler
Allow: /

User-agent: Cotoyogi
Allow: /

User-agent: Crawl4AI
Allow: /

User-agent: Crawlspace
Allow: /

User-agent: DataForSeoBot
Allow: /

User-agent: Diffbot
Allow: /

User-agent: dotbot
Allow: /

User-agent: DuckAssistBot
Allow: /

User-agent: DuckDuckBot
Allow: /

User-agent: EchoboxBot
Allow: /

User-agent: ExaSearchBot
Allow: /

User-agent: FacebookBot
Allow: /

User-agent: facebookexternalhit
Allow: /

User-agent: Factset_spyderbot
Allow: /

User-agent: FeedFetcher-Google   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: FirecrawlAgent
Allow: /

User-agent: Google-Agent   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-CloudVertexBot
Allow: /

User-agent: Google-CWS   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-Extended   # control token, no crawler uses this user-agent
Allow: /

User-agent: Google-GeminiNotebook   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-InspectionTool
Allow: /

User-agent: Google-Pinpoint   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-Read-Aloud   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-Safety   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Google-Site-Verification   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Googlebot
Allow: /

User-agent: Googlebot-Image
Allow: /

User-agent: Googlebot-News
Allow: /

User-agent: Googlebot-Video
Allow: /

User-agent: GoogleMessages   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: GoogleOther
Allow: /

User-agent: GoogleOther-Image
Allow: /

User-agent: GoogleOther-Video
Allow: /

User-agent: GoogleProducer   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: GPTBot
Allow: /

User-agent: ia_archiver
Allow: /

User-agent: ICC-Crawler
Allow: /

User-agent: ImagesiftBot
Allow: /

User-agent: img2dataset
Allow: /

User-agent: ISSCyberRiskCrawler   # compliance disputed; enforce at the edge
Allow: /

User-agent: Kagibot
Allow: /

User-agent: KlaviyoAIBot
Allow: /

User-agent: LAIONDownloader   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: Lightpanda
Allow: /

User-agent: Linguee Bot   # compliance disputed; enforce at the edge
Allow: /

User-agent: Mediapartners-Google
Allow: /

User-agent: meta-externalagent
Allow: /

User-agent: meta-externalfetcher
Allow: /

User-agent: Meta-WebIndexer
Allow: /

User-agent: MistralAI-User
Allow: /

User-agent: MJ12bot
Allow: /

User-agent: MojeekBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: omgili
Allow: /

User-agent: omgilibot
Allow: /

User-agent: panscient.com
Allow: /

User-agent: Perplexity-User   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: PetalBot
Allow: /

User-agent: PhindBot
Allow: /

User-agent: Pinterestbot
Allow: /

User-agent: Poseidon Research Crawler
Allow: /

User-agent: QualifiedBot
Allow: /

User-agent: QuillBot
Allow: /

User-agent: Qwantbot
Allow: /

User-agent: Qwantbot-news
Allow: /

User-agent: Reflectionbot
Allow: /

User-agent: rogerbot
Allow: /

User-agent: SBIntuitionsBot
Allow: /

User-agent: Scrapy
Allow: /

User-agent: Screaming Frog SEO Spider
Allow: /

User-agent: SemrushBot
Allow: /

User-agent: SemrushBot-BA
Allow: /

User-agent: SemrushBot-ESI
Allow: /

User-agent: SemrushBot-FT
Allow: /

User-agent: SemrushBot-OCOB
Allow: /

User-agent: SemrushBot-SI
Allow: /

User-agent: SemrushBot-SWA
Allow: /

User-agent: SEOkicks
Allow: /

User-agent: serpstatbot
Allow: /

User-agent: SeznamBot
Allow: /

User-agent: ShapBot
Allow: /

User-agent: Sidetrade indexer bot
Allow: /

User-agent: SiteAuditBot
Allow: /

User-agent: Slackbot
Allow: /

User-agent: Slackbot-LinkExpanding
Allow: /

User-agent: SplitSignalBot
Allow: /

User-agent: Storebot-Google
Allow: /

User-agent: TerraCotta
Allow: /

User-agent: Thinkbot   # compliance disputed; enforce at the edge
Allow: /

User-agent: TikTokSpider   # compliance disputed; enforce at the edge
Allow: /

User-agent: Timpibot
Allow: /

User-agent: VelenPublicWebCrawler
Allow: /

User-agent: Webzio-Extended
Allow: /

User-agent: wpbot
Allow: /

User-agent: YaK
Allow: /

User-agent: YandexAdditional
Allow: /

User-agent: YandexAdditionalBot
Allow: /

User-agent: YandexBlogs
Allow: /

User-agent: YandexBot
Allow: /

User-agent: YandexCalendar
Allow: /

User-agent: YandexComBot
Allow: /

User-agent: YandexDirect
Allow: /

User-agent: YandexFavicons
Allow: /

User-agent: YandexImages
Allow: /

User-agent: YandexMarket
Allow: /

User-agent: YandexMedia
Allow: /

User-agent: YandexMetrika   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: YandexMobileBot
Allow: /

User-agent: YandexRenderResourcesBot
Allow: /

User-agent: YandexScreenshotBot
Allow: /

User-agent: YandexVideo
Allow: /

User-agent: YandexWebmaster
Allow: /

User-agent: Yeti
Allow: /

User-agent: YouBot
Allow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml