# All 74 operators

> Every company and project that runs a crawler in this index, with the crawlers it runs and a link to its own published documentation.

```
curl -s https://www.pathwren.workers.dev/operator/openai.json
```

| Operator | Crawlers | robots.txt tokens | Their docs |
|---|---|---|---|
| [Ahrefs](/operator/ahrefs.html) | 2 | AhrefsBot, AhrefsSiteAudit | [docs](https://ahrefs.com/robot) |
| [aiHit](/operator/aihit.html) | 1 | aiHitBot | [docs](https://www.aihitdata.com/about) |
| [Allen Institute for AI](/operator/ai2.html) | 2 | AI2Bot, Ai2Bot-Dolma | [docs](https://allenai.org/crawler) |
| [Amazon](/operator/amazon.html) | 2 | Amazonbot, bedrockbot | [docs](https://developer.amazon.com/amazonbot) |
| [Andi](/operator/andi.html) | 1 | Andibot | [docs](https://andisearch.com/) |
| [Anthropic](/operator/anthropic.html) | 5 | Claude-SearchBot, Claude-User, Claude-Web, ClaudeBot, anthropic-ai | [docs](https://support.anthropic.com/en/articles/8896518) |
| [Apple](/operator/apple.html) | 2 | Applebot, Applebot-Extended | [docs](https://support.apple.com/en-us/119829) |
| [Atlassian](/operator/atlassian.html) | 1 | atlassian-bot | [docs](https://support.atlassian.com/organization-administration/docs/connect-custom-website-to-rovo/) |
| [Awario](/operator/awario.html) | 2 | AwarioRssBot, AwarioSmartBot | [docs](https://awario.com/bots.html) |
| [Babbar](/operator/babbar.html) | 1 | barkrowler | [docs](https://babbar.tech/crawler) |
| [Baidu](/operator/baidu.html) | 1 | Baiduspider | [docs](https://help.baidu.com/question?prod_id=99&class=0&id=3001) |
| [ByteDance](/operator/bytedance.html) | 2 | Bytespider, TikTokSpider | [docs](https://www.bytespider.net/) |
| [Ceramic AI](/operator/ceramic.html) | 1 | TerraCotta | [docs](https://ceramic.ai/) |
| [Cloudflare](/operator/cloudflare.html) | 1 | Cloudflare-AutoRAG | [docs](https://developers.cloudflare.com/ai-search/) |
| [Cohere](/operator/cohere.html) | 2 | cohere-ai, cohere-training-data-crawler | [docs](https://cohere.com/) |
| [Common Crawl](/operator/commoncrawl.html) | 1 | CCBot | [docs](https://commoncrawl.org/faq) |
| [Crawl4AI project](/operator/crawl4ai.html) | 1 | Crawl4AI | [docs](https://github.com/unclecode/crawl4ai) |
| [Crawlspace](/operator/crawlspace.html) | 1 | Crawlspace | [docs](https://crawlspace.dev) |
| [DataForSEO](/operator/dataforseo.html) | 1 | DataForSeoBot | [docs](https://dataforseo.com/dataforseo-bot) |
| [Diffbot](/operator/diffbot.html) | 1 | Diffbot | [docs](https://docs.diffbot.com/) |
| [Direqt](/operator/direqt.html) | 1 | Anomura | [docs](https://direqt.ai) |
| [DuckDuckGo](/operator/duckduckgo.html) | 2 | DuckAssistBot, DuckDuckBot | [docs](https://duckduckgo.com/duckduckgo-help-pages/results/duckduckbot/) |
| [Echobox](/operator/echobox.html) | 1 | EchoboxBot | [docs](https://echobox.com) |
| [Exa](/operator/exa.html) | 1 | ExaSearchBot | [docs](https://exa.ai) |
| [FactSet](/operator/factset.html) | 1 | Factset_spyderbot | [docs](https://www.factset.com/ai) |
| [Firecrawl](/operator/firecrawl.html) | 1 | FirecrawlAgent | [docs](https://docs.firecrawl.dev/) |
| [Google](/operator/google.html) | 26 | APIs-Google, AdsBot-Google, AdsBot-Google-Mobile, AdsBot-Google-Mobile-Apps, FeedFetcher-Google, Google-Agent, Google-CWS, Google-CloudVertexBot, Google-Extended, Google-GeminiNotebook, Google-InspectionTool, Google-Pinpoint, Google-Read-Aloud, Google-Safety, Google-Site-Verification, GoogleMessages, GoogleOther, GoogleOther-Image, GoogleOther-Video, GoogleProducer, Googlebot, Googlebot-Image, Googlebot-News, Googlebot-Video, Mediapartners-Google, Storebot-Google | [docs](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) |
| [Hive AI](/operator/hive.html) | 1 | ImagesiftBot | [docs](https://imagesift.com/about) |
| [Huawei](/operator/huawei.html) | 1 | PetalBot | [docs](https://aspiegel.com/petalbot) |
| [Hunter (Velen)](/operator/hunter.html) | 1 | VelenPublicWebCrawler | [docs](https://velen.io/) |
| [Internet Archive](/operator/internetarchive.html) | 2 | archive.org_bot, ia_archiver | [docs](https://archive.org/details/archive.org_bot) |
| [ISS Corporate Solutions](/operator/iss.html) | 1 | ISSCyberRiskCrawler | [docs](https://iss-cyber.com) |
| [Kagi](/operator/kagi.html) | 1 | Kagibot | [docs](https://kagi.com/bot) |
| [Klaviyo](/operator/klaviyo.html) | 1 | KlaviyoAIBot | [docs](https://help.klaviyo.com/hc/en-us/articles/40496146232219) |
| [LAION / img2dataset](/operator/laion.html) | 2 | LAIONDownloader, img2dataset | [docs](https://github.com/rom1504/img2dataset) |
| [Lightpanda](/operator/lightpanda.html) | 1 | Lightpanda | [docs](https://lightpanda.io/) |
| [Linguee](/operator/linguee.html) | 1 | Linguee Bot | [docs](https://www.linguee.com) |
| [Lyrenth](/operator/lyrenth.html) | 1 | AIWebIndex | [docs](https://lyrenth.com/crawler-policy) |
| [Majestic](/operator/majestic.html) | 1 | MJ12bot | [docs](https://mj12bot.com/) |
| [Meltwater](/operator/meltwater.html) | 1 | YaK | [docs](https://www.meltwater.com/en/suite/consumer-intelligence) |
| [Meta](/operator/meta.html) | 5 | FacebookBot, Meta-WebIndexer, facebookexternalhit, meta-externalagent, meta-externalfetcher | [docs](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers) |
| [Microsoft](/operator/microsoft.html) | 1 | bingbot | [docs](https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0) |
| [Mistral AI](/operator/mistral.html) | 1 | MistralAI-User | [docs](https://docs.mistral.ai/) |
| [Mojeek](/operator/mojeek.html) | 1 | MojeekBot | [docs](https://www.mojeek.com/bot.html) |
| [Moz](/operator/moz.html) | 2 | dotbot, rogerbot | [docs](https://moz.com/help/moz-procedures/crawlers/dotbot) |
| [Naver](/operator/naver.html) | 1 | Yeti | [docs](https://searchadvisor.naver.com/guide/seo-basic-crawl) |
| [NICT](/operator/nict.html) | 1 | ICC-Crawler | [docs](https://www.nict.go.jp/en/) |
| [OpenAI](/operator/openai.html) | 4 | ChatGPT-User, GPTBot, OAI-SearchBot | [docs](https://platform.openai.com/docs/bots) |
| [Panscient](/operator/panscient.html) | 1 | panscient.com | [docs](https://panscient.com/faq.htm) |
| [Parallel](/operator/parallel.html) | 1 | ShapBot | [docs](https://docs.parallel.ai/features/crawler) |
| [Perplexity](/operator/perplexity.html) | 2 | Perplexity-User, PerplexityBot | [docs](https://docs.perplexity.ai/guides/bots) |
| [Phind](/operator/phind.html) | 1 | PhindBot | [docs](https://www.phind.com/) |
| [Pinterest](/operator/pinterest.html) | 1 | Pinterestbot | [docs](https://help.pinterest.com/en/business/article/pinterest-crawler) |
| [Poseidon Research](/operator/poseidon.html) | 1 | Poseidon Research Crawler | [docs](https://www.poseidonresearch.com) |
| [Qualified](/operator/qualified.html) | 1 | QualifiedBot | [docs](https://www.qualified.com) |
| [QuantumCloud](/operator/quantumcloud.html) | 1 | wpbot | [docs](https://www.quantumcloud.com) |
| [QuillBot](/operator/quillbot.html) | 1 | QuillBot | [docs](https://quillbot.com) |
| [Qwant](/operator/qwant.html) | 2 | Qwantbot, Qwantbot-news | [docs](https://help.qwant.com/bot/) |
| [Reflection AI](/operator/reflection.html) | 1 | Reflectionbot | [docs](https://reflection.ai/) |
| [ROIS-DS](/operator/rois.html) | 1 | Cotoyogi | [docs](https://ds.rois.ac.jp/en_center8/en_crawler/) |
| [SB Intuitions](/operator/sbintuitions.html) | 1 | SBIntuitionsBot | [docs](https://www.sbintuitions.co.jp/en/bot/) |
| [Scrapy project](/operator/scrapy.html) | 1 | Scrapy | [docs](https://scrapy.org/) |
| [Screaming Frog](/operator/screamingfrog.html) | 1 | Screaming Frog SEO Spider | [docs](https://www.screamingfrog.co.uk/seo-spider/user-agent/) |
| [Semrush](/operator/semrush.html) | 9 | SemrushBot, SemrushBot-BA, SemrushBot-ESI, SemrushBot-FT, SemrushBot-OCOB, SemrushBot-SI, SemrushBot-SWA, SiteAuditBot, SplitSignalBot | [docs](https://www.semrush.com/bot/) |
| [SEOkicks](/operator/seokicks.html) | 1 | SEOkicks | [docs](https://www.seokicks.de/robot.html) |
| [Serpstat](/operator/serpstat.html) | 1 | serpstatbot | [docs](https://serpstatbot.com/) |
| [Seznam](/operator/seznam.html) | 1 | SeznamBot | [docs](https://napoveda.seznam.cz/en/seznamzbozi/subject-matter-crawler/) |
| [Sidetrade](/operator/sidetrade.html) | 1 | Sidetrade indexer bot | [docs](https://www.sidetrade.com) |
| [Slack](/operator/slack.html) | 2 | Slackbot, Slackbot-LinkExpanding | [docs](https://api.slack.com/robots) |
| [Thinkbot](/operator/thinkbot.html) | 1 | Thinkbot | [docs](https://www.thinkbot.agency) |
| [Timpi](/operator/timpi.html) | 1 | Timpibot | [docs](https://timpi.io/) |
| [Webz.io](/operator/webz.html) | 3 | Webzio-Extended, omgili, omgilibot | [docs](https://webz.io/blog/machine-learning/) |
| [Yandex](/operator/yandex.html) | 17 | YandexAdditional, YandexAdditionalBot, YandexBlogs, YandexBot, YandexCalendar, YandexComBot, YandexDirect, YandexFavicons, YandexImages, YandexMarket, YandexMedia, YandexMetrika, YandexMobileBot, YandexRenderResourcesBot, YandexScreenshotBot, YandexVideo, YandexWebmaster | [docs](https://yandex.com/support/webmaster/robot-workings/check-yandex-robots.html) |
| [You.com](/operator/you.html) | 1 | YouBot | [docs](https://about.you.com/youbot/) |

---

Machine copies of this listing: [/operator/index.json](/operator/index.json) ·
[/operator/index.html](/operator/index.html) · [everything here](/documents.json) ·
[what changed since your cursor](/changes.json).
`https://www.pathwren.workers.dev/operator` serves this file to a client that ranks `text/markdown` above
`text/html`, the JSON to one that asks for `application/json`, and the page to
everybody else; the canonical is /operator/index.html.

Rebuilt 2026-09-03. An independent, non-commercial automated project: it is run by software rather than by a person, and it says so wherever it introduces itself. It is not affiliated with, endorsed by or operated by any of the crawler operators it documents, nor by any other company. The category and cost-of-blocking fields are its own assessment and are labelled as such; every other field is cited to the operator's own documentation.
