Every AI crawler, and how to tell if it is real
This directory lists 150 user-agent tokens used by AI training crawlers, answer engines, assistants and autonomous agents. Each entry names the operator, gives the string as it appears in your logs, says whether the operator has committed to robots.txt, and states exactly how a request can be verified. Only 33 of them publish an IP range file, a reverse-DNS convention or an ASN you can check, so for the rest the user agent is a claim and nothing more.
Base facts come from the ai.robots.txt community list and from operator documentation. User-agent strings, IP range files and reverse-DNS conventions were checked against the operators' own published sources. Where nothing is published, the entry says so instead of guessing.
Block every AI crawler at once
One rule per token, covering every entry in this directory including the legacy aliases. Paste it at the root of your robots.txt. Bear in mind that 9 of these are documented as ignoring the file entirely, so pair it with edge rules if the traffic actually costs you money.
User-agent: AddSearchBot
User-agent: AgentTimes
User-agent: AI2Bot
User-agent: AI2Bot-DeepResearchEval
User-agent: Ai2Bot-Dolma
User-agent: aiHitBot
User-agent: AIWebIndex
User-agent: amazon-kendra
User-agent: amazon-QBusiness
User-agent: Amazonbot
User-agent: AmazonBuyForMe
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: Andibot
User-agent: Anomura
User-agent: anthropic-ai
User-agent: ApifyBot
User-agent: ApifyWebsiteContentCrawler
User-agent: Applebot
User-agent: Applebot-Extended
User-agent: Aranet-SearchBot
User-agent: atlassian-bot
User-agent: Awario
User-agent: AzureAI-SearchBot
User-agent: bedrockbot
User-agent: bigsur.ai
User-agent: Bingbot
User-agent: Bravebot
User-agent: Brightbot
User-agent: Brightbot 1.0
User-agent: BuddyBot
User-agent: Bytespider
User-agent: CCBot
User-agent: Channel3Bot
User-agent: ChatGLM-Spider
User-agent: ChatGPT Agent
User-agent: ChatGPT-User
User-agent: Claude-Code
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Claude-Web
User-agent: ClaudeBot
User-agent: Cloudflare-AutoRAG
User-agent: CloudVertexBot
User-agent: Code
User-agent: cohere-ai
User-agent: cohere-training-data-crawler
User-agent: Cotoyogi
User-agent: CragCrawler
User-agent: Crawl4AI
User-agent: Crawlspace
User-agent: Cursor
User-agent: Datenbank Crawler
User-agent: DeepSeekBot
User-agent: Devin
User-agent: Diffbot
User-agent: DuckAssistBot
User-agent: DuckDuckBot
User-agent: Echobot Bot
User-agent: EchoboxBot
User-agent: ExaBot
User-agent: FacebookBot
User-agent: facebookexternalhit
User-agent: Factset_spyderbot
User-agent: FirecrawlAgent
User-agent: FriendlyCrawler
User-agent: GeistHaus-PageFetcher
User-agent: Gemini-Deep-Research
User-agent: Google-Agent
User-agent: Google-CloudVertexBot
User-agent: Google-Extended
User-agent: Google-Firebase
User-agent: Google-Gemini-CLI
User-agent: Google-NotebookLM
User-agent: GoogleAgent-Mariner
User-agent: GoogleAgent-URLContext
User-agent: Googlebot
User-agent: GoogleOther
User-agent: GoogleOther-Image
User-agent: GoogleOther-Video
User-agent: GPTBot
User-agent: HenkBot
User-agent: iAskBot
User-agent: iaskspider
User-agent: iaskspider/2.0
User-agent: ICC-Crawler
User-agent: ImagesiftBot
User-agent: imageSpider
User-agent: img2dataset
User-agent: ISSCyberRiskCrawler
User-agent: kagi-fetcher
User-agent: Kangaroo Bot
User-agent: Kimi-User
User-agent: KlaviyoAIBot
User-agent: KunatoCrawler
User-agent: laion-huggingface-processor
User-agent: LAIONDownloader
User-agent: LCC
User-agent: LinerBot
User-agent: Linguee Bot
User-agent: LinkupBot
User-agent: Manus-User
User-agent: meta-externalagent
User-agent: Meta-ExternalAgent
User-agent: meta-externalfetcher
User-agent: Meta-ExternalFetcher
User-agent: meta-webindexer
User-agent: MistralAI-User
User-agent: MistralAI-User/1.0
User-agent: Mozilla-Tabstack
User-agent: MyCentralAIScraperBot
User-agent: NagetBot
User-agent: netEstate Imprint Crawler
User-agent: newsai
User-agent: NotebookLM
User-agent: NovaAct
User-agent: OAI-SearchBot
User-agent: omgili
User-agent: omgilibot
User-agent: OpenAI
User-agent: opencode
User-agent: Operator
User-agent: PanguBot
User-agent: Panscient
User-agent: panscient.com
User-agent: Perplexity-User
User-agent: PerplexityBot
User-agent: PetalBot
User-agent: PhindBot
User-agent: Poggio-Citations
User-agent: Poseidon Research Crawler
User-agent: QualifiedBot
User-agent: Querit-SearchBot
User-agent: QueritBot
User-agent: QuillBot
User-agent: quillbot.com
User-agent: SBIntuitionsBot
User-agent: Scrapy
User-agent: SemrushBot-OCOB
User-agent: SemrushBot-SWA
User-agent: Shap-User
User-agent: ShapBot
User-agent: Sidetrade indexer bot
User-agent: Spider
User-agent: TavilyBot
User-agent: Terra Cotta
User-agent: TerraCotta
User-agent: Thinkbot
User-agent: TikTokSpider
User-agent: Timpibot
User-agent: TongyiBot
User-agent: Trae
User-agent: TwinAgent
User-agent: UseAI
User-agent: VelenPublicWebCrawler
User-agent: WARDBot
User-agent: webzio-extended
User-agent: Webzio-Extended
User-agent: wpbot
User-agent: WRTNBot
User-agent: YaK
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
User-agent: YiyanBot
User-agent: YouBot
User-agent: ZanistaBot
Disallow: /User-agent: AI2Bot
User-agent: Ai2Bot-Dolma
User-agent: anthropic-ai
User-agent: Applebot-Extended
User-agent: Bytespider
User-agent: ClaudeBot
User-agent: cohere-ai
User-agent: Cotoyogi
User-agent: DeepSeekBot
User-agent: FacebookBot
User-agent: Factset_spyderbot
User-agent: FriendlyCrawler
User-agent: Google-Extended
User-agent: GPTBot
User-agent: ICC-Crawler
User-agent: img2dataset
User-agent: ISSCyberRiskCrawler
User-agent: Linguee Bot
User-agent: meta-externalagent
User-agent: Meta-ExternalAgent
User-agent: Scrapy
User-agent: Sidetrade indexer bot
User-agent: TikTokSpider
Disallow: /Keeps AI search crawlers, so you can still be cited in answers.
User-agent: AgentTimes
User-agent: aiHitBot
User-agent: AIWebIndex
User-agent: ApifyBot
User-agent: ApifyWebsiteContentCrawler
User-agent: Awario
User-agent: bedrockbot
User-agent: Bravebot
User-agent: Brightbot
User-agent: Brightbot 1.0
User-agent: CCBot
User-agent: ChatGLM-Spider
User-agent: cohere-training-data-crawler
User-agent: CragCrawler
User-agent: Crawlspace
User-agent: Datenbank Crawler
User-agent: Diffbot
User-agent: Echobot Bot
User-agent: ExaBot
User-agent: FirecrawlAgent
User-agent: HenkBot
User-agent: imageSpider
User-agent: Kangaroo Bot
User-agent: laion-huggingface-processor
User-agent: LAIONDownloader
User-agent: LCC
User-agent: Mozilla-Tabstack
User-agent: MyCentralAIScraperBot
User-agent: netEstate Imprint Crawler
User-agent: omgili
User-agent: omgilibot
User-agent: PanguBot
User-agent: Querit-SearchBot
User-agent: QueritBot
User-agent: SemrushBot-OCOB
User-agent: ShapBot
User-agent: Spider
User-agent: TavilyBot
User-agent: Terra Cotta
User-agent: TerraCotta
User-agent: VelenPublicWebCrawler
User-agent: WARDBot
User-agent: webzio-extended
User-agent: Webzio-Extended
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
Disallow: /Stops resale of your content while leaving the model vendors alone.
Training crawlers21
Collect pages that end up in a model's training corpus. Blocking these is the opt-out most publishers mean when they say they blocked AI.
AI search crawlers27
Build the index an AI answer engine cites from. Blocking one removes you from its answers, which is usually not what you want.
Assistant fetchers26
Fetch a single page because a person asked an assistant about it. Traffic is spiky, and several of these ignore robots.txt on purpose.
Autonomous agents17
Drive a browser through multi-step tasks. They look like a human session until you check the headers or the source IP.
Coding agents7
Pull docs and examples while writing code. The source IP is usually a developer laptop or a CI runner, not the vendor.
Data scrapers and resellers40
Crawl once, then sell or redistribute the result. Blocking stops resale rather than one model reading the page.
Other automated crawlers12
Bots that touch AI workflows without fitting the categories above.
Stop maintaining range files by hand
Every operator publishes its ranges in a different place, in a different format, and changes them without notice. Agentscan does the reverse-DNS and range matching for you and returns a verdict per request, so your edge can allow the real search crawlers and challenge everything wearing their names.