# AI Crawler Index — https://www.pathwren.workers.dev # Crawlers of every kind are welcome here, named explicitly and on purpose. # This site is a reference about you; there is nothing to withhold from you. # Full policy, and seven other ready-made ones: https://www.pathwren.workers.dev/policy/ User-agent: * Allow: / User-agent: AdsBot-Google Allow: / User-agent: AdsBot-Google-Mobile Allow: / User-agent: AdsBot-Google-Mobile-Apps Allow: / User-agent: AI2Bot Allow: / User-agent: Ai2Bot-Dolma Allow: / User-agent: aiHitBot Allow: / User-agent: AIWebIndex Allow: / User-agent: Amazonbot Allow: / User-agent: Andibot Allow: / User-agent: Anomura Allow: / User-agent: anthropic-ai Allow: / User-agent: APIs-Google Allow: / User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / User-agent: archive.org_bot Allow: / User-agent: atlassian-bot Allow: / User-agent: AwarioRssBot Allow: / User-agent: AwarioSmartBot Allow: / User-agent: Baiduspider Allow: / User-agent: bedrockbot Allow: / User-agent: bingbot Allow: / User-agent: Bytespider Allow: / User-agent: CCBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: Claude-Web Allow: / User-agent: ClaudeBot Allow: / User-agent: Cloudflare-AutoRAG Allow: / User-agent: cohere-ai Allow: / User-agent: cohere-training-data-crawler Allow: / User-agent: Cotoyogi Allow: / User-agent: Crawl4AI Allow: / User-agent: Crawlspace Allow: / User-agent: Diffbot Allow: / User-agent: DuckAssistBot Allow: / User-agent: DuckDuckBot Allow: / User-agent: EchoboxBot Allow: / User-agent: ExaSearchBot Allow: / User-agent: FacebookBot Allow: / User-agent: facebookexternalhit Allow: / User-agent: Factset_spyderbot Allow: / User-agent: FeedFetcher-Google Allow: / User-agent: FirecrawlAgent Allow: / User-agent: Google-Agent Allow: / User-agent: Google-CloudVertexBot Allow: / User-agent: Google-CWS Allow: / User-agent: Google-Extended Allow: / User-agent: Google-GeminiNotebook Allow: / User-agent: Google-InspectionTool Allow: / User-agent: Google-Pinpoint Allow: / User-agent: Google-Read-Aloud Allow: / User-agent: Google-Safety Allow: / User-agent: Google-Site-Verification Allow: / User-agent: Googlebot Allow: / User-agent: Googlebot-Image Allow: / User-agent: Googlebot-News Allow: / User-agent: Googlebot-Video Allow: / User-agent: GoogleMessages Allow: / User-agent: GoogleOther Allow: / User-agent: GoogleOther-Image Allow: / User-agent: GoogleOther-Video Allow: / User-agent: GoogleProducer Allow: / User-agent: GPTBot Allow: / User-agent: ia_archiver Allow: / User-agent: ICC-Crawler Allow: / User-agent: ImagesiftBot Allow: / User-agent: img2dataset Allow: / User-agent: ISSCyberRiskCrawler Allow: / User-agent: Kagibot Allow: / User-agent: KlaviyoAIBot Allow: / User-agent: LAIONDownloader Allow: / User-agent: Lightpanda Allow: / User-agent: Linguee Bot Allow: / User-agent: Mediapartners-Google Allow: / User-agent: meta-externalagent Allow: / User-agent: meta-externalfetcher Allow: / User-agent: Meta-WebIndexer Allow: / User-agent: MistralAI-User Allow: / User-agent: MojeekBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: omgili Allow: / User-agent: omgilibot Allow: / User-agent: panscient.com Allow: / User-agent: Perplexity-User Allow: / User-agent: PerplexityBot Allow: / User-agent: PetalBot Allow: / User-agent: PhindBot Allow: / User-agent: Pinterestbot Allow: / User-agent: Poseidon Research Crawler Allow: / User-agent: QualifiedBot Allow: / User-agent: QuillBot Allow: / User-agent: Qwantbot Allow: / User-agent: Qwantbot-news Allow: / User-agent: Reflectionbot Allow: / User-agent: SBIntuitionsBot Allow: / User-agent: Scrapy Allow: / User-agent: Screaming Frog SEO Spider Allow: / User-agent: SemrushBot-OCOB Allow: / User-agent: SeznamBot Allow: / User-agent: ShapBot Allow: / User-agent: Sidetrade indexer bot Allow: / User-agent: Slackbot Allow: / User-agent: Slackbot-LinkExpanding Allow: / User-agent: Storebot-Google Allow: / User-agent: TerraCotta Allow: / User-agent: Thinkbot Allow: / User-agent: TikTokSpider Allow: / User-agent: Timpibot Allow: / User-agent: VelenPublicWebCrawler Allow: / User-agent: Webzio-Extended Allow: / User-agent: wpbot Allow: / User-agent: YaK Allow: / User-agent: YandexAdditional Allow: / User-agent: YandexAdditionalBot Allow: / User-agent: YandexBlogs Allow: / User-agent: YandexBot Allow: / User-agent: YandexCalendar Allow: / User-agent: YandexComBot Allow: / User-agent: YandexDirect Allow: / User-agent: YandexFavicons Allow: / User-agent: YandexImages Allow: / User-agent: YandexMarket Allow: / User-agent: YandexMedia Allow: / User-agent: YandexMetrika Allow: / User-agent: YandexMobileBot Allow: / User-agent: YandexRenderResourcesBot Allow: / User-agent: YandexScreenshotBot Allow: / User-agent: YandexVideo Allow: / User-agent: YandexWebmaster Allow: / User-agent: Yeti Allow: / User-agent: YouBot Allow: / # No crawl-delay: every page here is a small static file. Sitemap: https://www.pathwren.workers.dev/sitemap.xml # Agentic Resource Discovery (ARD, June 2026): the manifest of what an agent # can call here. Registries crawl robots.txt for this line — it is one of the # four paths their publisher guide asks for, and the one a crawler that never # guesses /.well-known/ depends on. Agentmap: https://www.pathwren.workers.dev/.well-known/ard.json