AI crawler directory

Every AI crawler, and how to tell if it is real

This directory lists 150 user-agent tokens used by AI training crawlers, answer engines, assistants and autonomous agents. Each entry names the operator, gives the string as it appears in your logs, says whether the operator has committed to robots.txt, and states exactly how a request can be verified. Only 33 of them publish an IP range file, a reverse-DNS convention or an ASN you can check, so for the rest the user agent is a claim and nothing more.

150
crawler tokens tracked
33
that publish a way to verify them
9
documented as ignoring robots.txt
88
with no published robots.txt policy

Base facts come from the ai.robots.txt community list and from operator documentation. User-agent strings, IP range files and reverse-DNS conventions were checked against the operators' own published sources. Where nothing is published, the entry says so instead of guessing.

robots.txt

Block every AI crawler at once

One rule per token, covering every entry in this directory including the legacy aliases. Paste it at the root of your robots.txt. Bear in mind that 9 of these are documented as ignoring the file entirely, so pair it with edge rules if the traffic actually costs you money.

Block all 150 crawlers
User-agent: AddSearchBot
User-agent: AgentTimes
User-agent: AI2Bot
User-agent: AI2Bot-DeepResearchEval
User-agent: Ai2Bot-Dolma
User-agent: aiHitBot
User-agent: AIWebIndex
User-agent: amazon-kendra
User-agent: amazon-QBusiness
User-agent: Amazonbot
User-agent: AmazonBuyForMe
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: Andibot
User-agent: Anomura
User-agent: anthropic-ai
User-agent: ApifyBot
User-agent: ApifyWebsiteContentCrawler
User-agent: Applebot
User-agent: Applebot-Extended
User-agent: Aranet-SearchBot
User-agent: atlassian-bot
User-agent: Awario
User-agent: AzureAI-SearchBot
User-agent: bedrockbot
User-agent: bigsur.ai
User-agent: Bingbot
User-agent: Bravebot
User-agent: Brightbot
User-agent: Brightbot 1.0
User-agent: BuddyBot
User-agent: Bytespider
User-agent: CCBot
User-agent: Channel3Bot
User-agent: ChatGLM-Spider
User-agent: ChatGPT Agent
User-agent: ChatGPT-User
User-agent: Claude-Code
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Claude-Web
User-agent: ClaudeBot
User-agent: Cloudflare-AutoRAG
User-agent: CloudVertexBot
User-agent: Code
User-agent: cohere-ai
User-agent: cohere-training-data-crawler
User-agent: Cotoyogi
User-agent: CragCrawler
User-agent: Crawl4AI
User-agent: Crawlspace
User-agent: Cursor
User-agent: Datenbank Crawler
User-agent: DeepSeekBot
User-agent: Devin
User-agent: Diffbot
User-agent: DuckAssistBot
User-agent: DuckDuckBot
User-agent: Echobot Bot
User-agent: EchoboxBot
User-agent: ExaBot
User-agent: FacebookBot
User-agent: facebookexternalhit
User-agent: Factset_spyderbot
User-agent: FirecrawlAgent
User-agent: FriendlyCrawler
User-agent: GeistHaus-PageFetcher
User-agent: Gemini-Deep-Research
User-agent: Google-Agent
User-agent: Google-CloudVertexBot
User-agent: Google-Extended
User-agent: Google-Firebase
User-agent: Google-Gemini-CLI
User-agent: Google-NotebookLM
User-agent: GoogleAgent-Mariner
User-agent: GoogleAgent-URLContext
User-agent: Googlebot
User-agent: GoogleOther
User-agent: GoogleOther-Image
User-agent: GoogleOther-Video
User-agent: GPTBot
User-agent: HenkBot
User-agent: iAskBot
User-agent: iaskspider
User-agent: iaskspider/2.0
User-agent: ICC-Crawler
User-agent: ImagesiftBot
User-agent: imageSpider
User-agent: img2dataset
User-agent: ISSCyberRiskCrawler
User-agent: kagi-fetcher
User-agent: Kangaroo Bot
User-agent: Kimi-User
User-agent: KlaviyoAIBot
User-agent: KunatoCrawler
User-agent: laion-huggingface-processor
User-agent: LAIONDownloader
User-agent: LCC
User-agent: LinerBot
User-agent: Linguee Bot
User-agent: LinkupBot
User-agent: Manus-User
User-agent: meta-externalagent
User-agent: Meta-ExternalAgent
User-agent: meta-externalfetcher
User-agent: Meta-ExternalFetcher
User-agent: meta-webindexer
User-agent: MistralAI-User
User-agent: MistralAI-User/1.0
User-agent: Mozilla-Tabstack
User-agent: MyCentralAIScraperBot
User-agent: NagetBot
User-agent: netEstate Imprint Crawler
User-agent: newsai
User-agent: NotebookLM
User-agent: NovaAct
User-agent: OAI-SearchBot
User-agent: omgili
User-agent: omgilibot
User-agent: OpenAI
User-agent: opencode
User-agent: Operator
User-agent: PanguBot
User-agent: Panscient
User-agent: panscient.com
User-agent: Perplexity-User
User-agent: PerplexityBot
User-agent: PetalBot
User-agent: PhindBot
User-agent: Poggio-Citations
User-agent: Poseidon Research Crawler
User-agent: QualifiedBot
User-agent: Querit-SearchBot
User-agent: QueritBot
User-agent: QuillBot
User-agent: quillbot.com
User-agent: SBIntuitionsBot
User-agent: Scrapy
User-agent: SemrushBot-OCOB
User-agent: SemrushBot-SWA
User-agent: Shap-User
User-agent: ShapBot
User-agent: Sidetrade indexer bot
User-agent: Spider
User-agent: TavilyBot
User-agent: Terra Cotta
User-agent: TerraCotta
User-agent: Thinkbot
User-agent: TikTokSpider
User-agent: Timpibot
User-agent: TongyiBot
User-agent: Trae
User-agent: TwinAgent
User-agent: UseAI
User-agent: VelenPublicWebCrawler
User-agent: WARDBot
User-agent: webzio-extended
User-agent: Webzio-Extended
User-agent: wpbot
User-agent: WRTNBot
User-agent: YaK
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
User-agent: YiyanBot
User-agent: YouBot
User-agent: ZanistaBot
Disallow: /
Block training crawlers only
User-agent: AI2Bot
User-agent: Ai2Bot-Dolma
User-agent: anthropic-ai
User-agent: Applebot-Extended
User-agent: Bytespider
User-agent: ClaudeBot
User-agent: cohere-ai
User-agent: Cotoyogi
User-agent: DeepSeekBot
User-agent: FacebookBot
User-agent: Factset_spyderbot
User-agent: FriendlyCrawler
User-agent: Google-Extended
User-agent: GPTBot
User-agent: ICC-Crawler
User-agent: img2dataset
User-agent: ISSCyberRiskCrawler
User-agent: Linguee Bot
User-agent: meta-externalagent
User-agent: Meta-ExternalAgent
User-agent: Scrapy
User-agent: Sidetrade indexer bot
User-agent: TikTokSpider
Disallow: /

Keeps AI search crawlers, so you can still be cited in answers.

Block data scrapers and resellers only
User-agent: AgentTimes
User-agent: aiHitBot
User-agent: AIWebIndex
User-agent: ApifyBot
User-agent: ApifyWebsiteContentCrawler
User-agent: Awario
User-agent: bedrockbot
User-agent: Bravebot
User-agent: Brightbot
User-agent: Brightbot 1.0
User-agent: CCBot
User-agent: ChatGLM-Spider
User-agent: cohere-training-data-crawler
User-agent: CragCrawler
User-agent: Crawlspace
User-agent: Datenbank Crawler
User-agent: Diffbot
User-agent: Echobot Bot
User-agent: ExaBot
User-agent: FirecrawlAgent
User-agent: HenkBot
User-agent: imageSpider
User-agent: Kangaroo Bot
User-agent: laion-huggingface-processor
User-agent: LAIONDownloader
User-agent: LCC
User-agent: Mozilla-Tabstack
User-agent: MyCentralAIScraperBot
User-agent: netEstate Imprint Crawler
User-agent: omgili
User-agent: omgilibot
User-agent: PanguBot
User-agent: Querit-SearchBot
User-agent: QueritBot
User-agent: SemrushBot-OCOB
User-agent: ShapBot
User-agent: Spider
User-agent: TavilyBot
User-agent: Terra Cotta
User-agent: TerraCotta
User-agent: VelenPublicWebCrawler
User-agent: WARDBot
User-agent: webzio-extended
User-agent: Webzio-Extended
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
Disallow: /

Stops resale of your content while leaving the model vendors alone.

Assistant fetchers26

Fetch a single page because a person asked an assistant about it. Traffic is spiky, and several of these ignore robots.txt on purpose.

CrawlerOperatorrobots.txtVerification
AI2Bot-DeepResearchEvalAi2, a non-profit AI research instituteNot documentedNone published, user agent onlyamazon-QBusinessAmazon Web ServicesNot documentedNone published, user agent onlyAmzn-UserAmazonNot documentedReverse DNS, forward confirmedbigsur.aiUnidentifiedNot documentedNone published, user agent onlyChatGPT-UserOpenAIYes, documentedPublished IP range fileClaude-UserAnthropicYes, documentedPublished IP range fileDuckAssistBotUnidentifiedYes, documentedPublished IP range fileGeistHaus-PageFetcherUnidentifiedNot documentedNone published, user agent onlyGemini-Deep-ResearchGoogleNot documentedNone published, user agent onlyGoogle-NotebookLMNotebookLMGoogleNot documentedPublished IP range fileGoogleAgent-URLContextGoogleNot documentedNone published, user agent onlykagi-fetcherUnidentifiedNot documentedNone published, user agent onlyKimi-UserMoonshot AINot documentedNone published, user agent onlyKlaviyoAIBotKlaviyoYes, documentedNone published, user agent onlyLinerBotUnidentifiedNot documentedNone published, user agent onlyMeta-ExternalFetchermeta-externalfetcherMetaNoOperator ASN checkMistralAI-UserMistralAI-User/1.0MistralNot documentedNone published, user agent onlyPerplexity-UserPerplexityNoPublished IP range filePhindBotphindNot documentedNone published, user agent onlyPoggio-CitationsUnidentifiedNot documentedNone published, user agent onlyQualifiedBotQualifiedNot documentedNone published, user agent onlySemrushBot-SWASemrushYes, documentedNone published, user agent onlyShap-UserParallelNot documentedNone published, user agent onlyTongyiBotAlibabaNot documentedNone published, user agent onlyUseAIUnidentifiedNot documentedNone published, user agent onlyYiyanBotBaidu that fetches web content for the yiyanNot documentedNone published, user agent only

Data scrapers and resellers40

Crawl once, then sell or redistribute the result. Blocking stops resale rather than one model reading the page.

CrawlerOperatorrobots.txtVerification
AgentTimesThe Agent TimesNot documentedNone published, user agent onlyaiHitBotaiHitYes, documentedNone published, user agent onlyAIWebIndexUnidentifiedNot documentedNone published, user agent onlyApifyBotUnidentifiedNot documentedNone published, user agent onlyApifyWebsiteContentCrawlerUnidentifiedNot documentedNone published, user agent onlyAwarioAwarioNot documentedNone published, user agent onlybedrockbotAmazonYes, documentedNone published, user agent onlyBravebotBraveYes, documentedNone published, user agent onlyBrightbotBrightbot 1.0UnidentifiedNot documentedNone published, user agent onlyCCBotCommon Crawl FoundationYes, documentedNone published, user agent onlyChatGLM-SpiderZhipu AINot documentedNone published, user agent onlycohere-training-data-crawlerCohereNot documentedNone published, user agent onlyCragCrawlerUnidentifiedNot documentedNone published, user agent onlyCrawlspaceCrawlspaceYes, documentedNone published, user agent onlyDatenbank CrawlerDatenbankNot documentedNone published, user agent onlyDiffbotDiffbotNot documentedNone published, user agent onlyEchobot BotEchoboxNot documentedNone published, user agent onlyExaBotExaNot documentedNone published, user agent onlyFirecrawlAgentFirecrawlYes, documentedNone published, user agent onlyHenkBotUnidentifiedNot documentedNone published, user agent onlyimageSpiderUnidentifiedNot documentedNone published, user agent onlyKangaroo BotKangaroo LLMNot documentedNone published, user agent onlylaion-huggingface-processorLAIONNot documentedNone published, user agent onlyLAIONDownloaderLarge-scale Artificial Intelligence Open NetworkNoNone published, user agent onlyLCCUnidentifiedNot documentedNone published, user agent onlyMozilla-TabstackMozillaYes, documentedNone published, user agent onlyMyCentralAIScraperBotUnidentifiedNot documentedNone published, user agent onlynetEstate Imprint CrawlernetEstateNot documentedNone published, user agent onlyomgiliomgilibotWebz.ioYes, documentedNone published, user agent onlyPanguBotthe Chinese company HuaweiNot documentedNone published, user agent onlyQueritBotQuerit-SearchBotQueritNot documentedNone published, user agent onlySemrushBot-OCOBSemrushYes, documentedNone published, user agent onlyShapBotParallelYes, documentedNone published, user agent onlySpiderUnidentifiedNot documentedNone published, user agent onlyTavilyBotTavilyNot documentedNone published, user agent onlyTerraCottaTerra CottaCeramic AIYes, documentedNone published, user agent onlyVelenPublicWebCrawlerVelen CrawlerYes, documentedNone published, user agent onlyWARDBotWEBSPARKNot documentedNone published, user agent onlyWebzio-Extendedwebzio-extendedUnidentifiedNot documentedNone published, user agent onlyYandexAdditionalYandexAdditionalBotYandexYes, documentedReverse DNS, forward confirmed
Agentscan

Stop maintaining range files by hand

Every operator publishes its ranges in a different place, in a different format, and changes them without notice. Agentscan does the reverse-DNS and range matching for you and returns a verdict per request, so your edge can allow the real search crawlers and challenge everything wearing their names.