AI crawler directory

Every AI crawler, and how to tell if it is real

119 of the 153 AI crawler tokens in this directory publish no way to prove a request is really theirs: no IP range file, no reverse-DNS convention, no operator ASN. For those 78%, the user agent is a claim and nothing more, and anyone can type it. The remaining 34 can be checked, and each entry below says exactly how: who runs the crawler, the string as it appears in your logs, whether the operator has committed to robots.txt, and what evidence exists either way.

78%
publish no way to prove a request is really theirs
153
crawler tokens tracked in total
9
documented as ignoring robots.txt
90
with no published robots.txt policy

Base facts come from the ai.robots.txt community list and from operator documentation. User-agent strings, IP range files and reverse-DNS conventions were checked against the operators' own published sources. Where nothing is published, the entry says so instead of guessing.

Last updated

robots.txt

Block every AI crawler at once

One rule per token, covering every entry in this directory including the legacy aliases. Paste it at the root of your robots.txt. Bear in mind that 9 of these are documented as ignoring the file entirely, so pair it with edge rules if the traffic actually costs you money.

Block all 153 crawlers
User-agent: AddSearchBot
User-agent: AgentTimes
User-agent: AI2Bot
User-agent: AI2Bot-DeepResearchEval
User-agent: Ai2Bot-Dolma
User-agent: aiHitBot
User-agent: AIWebIndex
User-agent: AIWebIndex-Agent
User-agent: amazon-kendra
User-agent: amazon-QBusiness
User-agent: Amazonbot
User-agent: AmazonBuyForMe
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: Andibot
User-agent: Anomura
User-agent: anthropic-ai
User-agent: ApifyBot
User-agent: ApifyWebsiteContentCrawler
User-agent: Applebot
User-agent: Applebot-Extended
User-agent: Aranet-SearchBot
User-agent: atlassian-bot
User-agent: Awario
User-agent: AzureAI-SearchBot
User-agent: bedrockbot
User-agent: bigsur.ai
User-agent: Bingbot
User-agent: Bravebot
User-agent: Brightbot
User-agent: Brightbot 1.0
User-agent: BuddyBot
User-agent: Bytespider
User-agent: CCBot
User-agent: Channel3Bot
User-agent: ChatGLM-Spider
User-agent: ChatGPT Agent
User-agent: ChatGPT-User
User-agent: Claude-Code
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Claude-Web
User-agent: ClaudeBot
User-agent: Cloudflare-AutoRAG
User-agent: CloudVertexBot
User-agent: Code
User-agent: cohere-ai
User-agent: cohere-training-data-crawler
User-agent: Cotoyogi
User-agent: CragCrawler
User-agent: Crawl4AI
User-agent: Crawlspace
User-agent: Cursor
User-agent: Datenbank Crawler
User-agent: DeepSeekBot
User-agent: Devin
User-agent: Diffbot
User-agent: DuckAssistBot
User-agent: DuckDuckBot
User-agent: Echobot Bot
User-agent: EchoboxBot
User-agent: ExaBot
User-agent: ExaSearchBot
User-agent: FacebookBot
User-agent: facebookexternalhit
User-agent: Factset_spyderbot
User-agent: FirecrawlAgent
User-agent: FriendlyCrawler
User-agent: GeistHaus-PageFetcher
User-agent: Gemini-Deep-Research
User-agent: Google-Agent
User-agent: Google-CloudVertexBot
User-agent: Google-Extended
User-agent: Google-Firebase
User-agent: Google-Gemini-CLI
User-agent: Google-NotebookLM
User-agent: GoogleAgent-Mariner
User-agent: GoogleAgent-URLContext
User-agent: Googlebot
User-agent: GoogleOther
User-agent: GoogleOther-Image
User-agent: GoogleOther-Video
User-agent: GPTBot
User-agent: HenkBot
User-agent: iAskBot
User-agent: iaskspider
User-agent: iaskspider/2.0
User-agent: ICC-Crawler
User-agent: ImagesiftBot
User-agent: imageSpider
User-agent: img2dataset
User-agent: ISSCyberRiskCrawler
User-agent: kagi-fetcher
User-agent: Kangaroo Bot
User-agent: Kimi-User
User-agent: KlaviyoAIBot
User-agent: KunatoCrawler
User-agent: laion-huggingface-processor
User-agent: LAIONDownloader
User-agent: LCC
User-agent: Lightpanda
User-agent: LinerBot
User-agent: Linguee Bot
User-agent: LinkupBot
User-agent: Manus-User
User-agent: meta-externalagent
User-agent: Meta-ExternalAgent
User-agent: meta-externalfetcher
User-agent: Meta-ExternalFetcher
User-agent: meta-webindexer
User-agent: MistralAI-User
User-agent: MistralAI-User/1.0
User-agent: Mozilla-Tabstack
User-agent: MyCentralAIScraperBot
User-agent: NagetBot
User-agent: netEstate Imprint Crawler
User-agent: newsai
User-agent: NotebookLM
User-agent: NovaAct
User-agent: OAI-SearchBot
User-agent: omgili
User-agent: omgilibot
User-agent: OpenAI
User-agent: opencode
User-agent: Operator
User-agent: PanguBot
User-agent: Panscient
User-agent: panscient.com
User-agent: Perplexity-User
User-agent: PerplexityBot
User-agent: PetalBot
User-agent: PhindBot
User-agent: Poggio-Citations
User-agent: Poseidon Research Crawler
User-agent: QualifiedBot
User-agent: Querit-SearchBot
User-agent: QueritBot
User-agent: QuillBot
User-agent: quillbot.com
User-agent: Reflectionbot
User-agent: SBIntuitionsBot
User-agent: Scrapy
User-agent: SemrushBot-OCOB
User-agent: SemrushBot-SWA
User-agent: Shap-User
User-agent: ShapBot
User-agent: Sidetrade indexer bot
User-agent: Spider
User-agent: TavilyBot
User-agent: Terra Cotta
User-agent: TerraCotta
User-agent: Thinkbot
User-agent: TikTokSpider
User-agent: Timpibot
User-agent: TongyiBot
User-agent: Trae
User-agent: TwinAgent
User-agent: UseAI
User-agent: VelenPublicWebCrawler
User-agent: WARDBot
User-agent: webzio-extended
User-agent: Webzio-Extended
User-agent: wpbot
User-agent: WRTNBot
User-agent: YaK
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
User-agent: YiyanBot
User-agent: YouBot
User-agent: ZanistaBot
Disallow: /
Block training crawlers only
User-agent: AI2Bot
User-agent: Ai2Bot-Dolma
User-agent: anthropic-ai
User-agent: Applebot-Extended
User-agent: Bytespider
User-agent: ClaudeBot
User-agent: cohere-ai
User-agent: Cotoyogi
User-agent: DeepSeekBot
User-agent: FacebookBot
User-agent: Factset_spyderbot
User-agent: FriendlyCrawler
User-agent: Google-Extended
User-agent: GPTBot
User-agent: ICC-Crawler
User-agent: img2dataset
User-agent: ISSCyberRiskCrawler
User-agent: Linguee Bot
User-agent: meta-externalagent
User-agent: Meta-ExternalAgent
User-agent: Scrapy
User-agent: Sidetrade indexer bot
User-agent: TikTokSpider
Disallow: /

Keeps AI search crawlers, so you can still be cited in answers.

Block dataset builders only
User-agent: AgentTimes
User-agent: aiHitBot
User-agent: ApifyBot
User-agent: ApifyWebsiteContentCrawler
User-agent: Awario
User-agent: bedrockbot
User-agent: Bravebot
User-agent: Brightbot
User-agent: Brightbot 1.0
User-agent: CCBot
User-agent: ChatGLM-Spider
User-agent: cohere-training-data-crawler
User-agent: CragCrawler
User-agent: Crawlspace
User-agent: Datenbank Crawler
User-agent: Diffbot
User-agent: Echobot Bot
User-agent: ExaBot
User-agent: FirecrawlAgent
User-agent: HenkBot
User-agent: imageSpider
User-agent: Kangaroo Bot
User-agent: laion-huggingface-processor
User-agent: LAIONDownloader
User-agent: LCC
User-agent: Lightpanda
User-agent: Mozilla-Tabstack
User-agent: MyCentralAIScraperBot
User-agent: netEstate Imprint Crawler
User-agent: omgili
User-agent: omgilibot
User-agent: PanguBot
User-agent: Querit-SearchBot
User-agent: QueritBot
User-agent: SemrushBot-OCOB
User-agent: ShapBot
User-agent: Spider
User-agent: TavilyBot
User-agent: Terra Cotta
User-agent: TerraCotta
User-agent: VelenPublicWebCrawler
User-agent: WARDBot
User-agent: webzio-extended
User-agent: Webzio-Extended
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
Disallow: /

Keeps your pages out of collected datasets while leaving the model vendors alone.

Training crawlers21

Collect pages that end up in a model's training corpus. Blocking these is the opt-out most publishers mean when they say they blocked AI.

Training crawlers: user-agent token, operator, robots.txt policy and how each one can be verified
CrawlerOperatorrobots.txtVerificationDetails
AI2BotAi2Bot-DolmaAi2Yes, documentedNone published, user agent only
anthropic-aiAnthropicNot documentedPublished IP range file
Applebot-ExtendedAppleYes, documentedNot applicable, robots.txt token only
BytespiderByteDanceNoNone published, user agent only
ClaudeBotAnthropicYes, documentedPublished IP range file
cohere-aiCohereNot documentedNone published, user agent only
CotoyogiROISYes, documentedNone published, user agent only
DeepSeekBotDeepSeekNoNone published, user agent only
FacebookBotMeta/FacebookYes, documentedOperator ASN check
Factset_spyderbotFactsetNot documentedNone published, user agent only
FriendlyCrawlerUnknownYes, documentedNone published, user agent only
Google-ExtendedGoogleYes, documentedNot applicable, robots.txt token only
GPTBotOpenAIYes, documentedPublished IP range file
ICC-CrawlerNICTYes, documentedNone published, user agent only
img2datasetimg2datasetNot documentedNone published, user agent only
ISSCyberRiskCrawlerISS-CorporateNoNone published, user agent only
Linguee BotLingueeNoNone published, user agent only
Meta-ExternalAgentmeta-externalagentMetaYes, documentedOperator ASN check
ScrapyZyteNot documentedNone published, user agent only
Sidetrade indexer botSidetradeNot documentedNone published, user agent only
TikTokSpiderByteDanceNot documentedNone published, user agent only

Assistant fetchers26

Fetch a single page because a person asked an assistant about it. Traffic is spiky, and some operators document that these user-requested fetches skip robots.txt.

Assistant fetchers: user-agent token, operator, robots.txt policy and how each one can be verified
CrawlerOperatorrobots.txtVerificationDetails
AI2Bot-DeepResearchEvalAi2, a non-profit AI research instituteNot documentedNone published, user agent only
amazon-QBusinessAmazon Web ServicesNot documentedNone published, user agent only
Amzn-UserAmazonNot documentedReverse DNS, forward confirmed
bigsur.aiUnidentifiedNot documentedNone published, user agent only
ChatGPT-UserOpenAIYes, documentedPublished IP range file
Claude-UserAnthropicYes, documentedPublished IP range file
DuckAssistBotUnidentifiedYes, documentedPublished IP range file
GeistHaus-PageFetcherUnidentifiedNot documentedNone published, user agent only
Gemini-Deep-ResearchGoogleNot documentedNone published, user agent only
Google-NotebookLMNotebookLMGoogleNot documentedPublished IP range file
GoogleAgent-URLContextGoogleNot documentedNone published, user agent only
kagi-fetcherUnidentifiedNot documentedNone published, user agent only
Kimi-UserMoonshot AINot documentedNone published, user agent only
KlaviyoAIBotKlaviyoYes, documentedNone published, user agent only
LinerBotUnidentifiedNot documentedNone published, user agent only
Meta-ExternalFetchermeta-externalfetcherMetaNoOperator ASN check
MistralAI-UserMistralAI-User/1.0MistralNot documentedNone published, user agent only
Perplexity-UserPerplexityNoPublished IP range file
PhindBotphindNot documentedNone published, user agent only
Poggio-CitationsUnidentifiedNot documentedNone published, user agent only
QualifiedBotQualifiedNot documentedNone published, user agent only
SemrushBot-SWASemrushYes, documentedNone published, user agent only
Shap-UserParallelNot documentedNone published, user agent only
TongyiBotAlibabaNot documentedNone published, user agent only
UseAIUnidentifiedNot documentedNone published, user agent only
YiyanBotBaidu that fetches web content for the yiyanNot documentedNone published, user agent only

Autonomous agents18

Drive a browser through multi-step tasks. They look like a human session until you check the headers or the source IP.

Autonomous agents: user-agent token, operator, robots.txt policy and how each one can be verified
CrawlerOperatorrobots.txtVerificationDetails
AmazonBuyForMeAmazonNot documentedNone published, user agent only
Aranet-SearchBotUnidentifiedNot documentedNone published, user agent only
ChatGPT AgentOpenAIYes, documentedPublished IP range file
Claude-WebAnthropicNot documentedPublished IP range file
Crawl4AICrawl4AI (open source)Not documentedNone published, user agent only
Google-AgentGoogleYes, documentedNone published, user agent only
Google-CloudVertexBotCloudVertexBotGoogleYes, documentedPublished IP range file
GoogleAgent-MarinerGoogleNot documentedNone published, user agent only
iAskBotiaskspider, iaskspider/2.0iAskNot documentedNone published, user agent only
KunatoCrawlerUnidentifiedNot documentedNone published, user agent only
Manus-UserButterfly Effect, a company based in ChinaNot documentedNone published, user agent only
NagetBotUnidentifiedNot documentedNone published, user agent only
newsaiUnidentifiedNot documentedNone published, user agent only
NovaActAmazonNot documentedNone published, user agent only
OperatorOpenAINot documentedPublished IP range file
ReflectionbotReflectionNot documentedNone published, user agent only
TwinAgentUnidentifiedNot documentedNone published, user agent only
WRTNBotWrtn TechnologiesNot documentedNone published, user agent only

Coding agents7

Pull docs and examples while writing code. The source IP is usually a developer laptop or a CI runner, not the vendor.

Coding agents: user-agent token, operator, robots.txt policy and how each one can be verified
CrawlerOperatorrobots.txtVerificationDetails
Claude-CodeAnthropicNot documentedPublished IP range file
CodeGitHub CopilotNot documentedNone published, user agent only
CursorCursorNot documentedNone published, user agent only
DevinDevin AIYes, documentedNone published, user agent only
Google-Gemini-CLIGoogleNot documentedNone published, user agent only
opencodeopencode (open source)Not documentedNone published, user agent only
TraeByteDanceNot documentedNone published, user agent only

Dataset builders40

Collect pages into datasets that may be reused, licensed or redistributed. Blocking limits future collection rather than one model reading the page.

Dataset builders: user-agent token, operator, robots.txt policy and how each one can be verified
CrawlerOperatorrobots.txtVerificationDetails
AgentTimesThe Agent TimesNot documentedNone published, user agent only
aiHitBotaiHitYes, documentedNone published, user agent only
ApifyBotUnidentifiedNot documentedNone published, user agent only
ApifyWebsiteContentCrawlerUnidentifiedNot documentedNone published, user agent only
AwarioAwarioNot documentedNone published, user agent only
bedrockbotAmazonYes, documentedNone published, user agent only
BravebotBraveYes, documentedNone published, user agent only
BrightbotBrightbot 1.0UnidentifiedNot documentedNone published, user agent only
CCBotCommon Crawl FoundationYes, documentedNone published, user agent only
ChatGLM-SpiderZhipu AINot documentedNone published, user agent only
cohere-training-data-crawlerCohereNot documentedNone published, user agent only
CragCrawlerUnidentifiedNot documentedNone published, user agent only
CrawlspaceCrawlspaceYes, documentedNone published, user agent only
Datenbank CrawlerDatenbankNot documentedNone published, user agent only
DiffbotDiffbotNot documentedNone published, user agent only
Echobot BotEchoboxNot documentedNone published, user agent only
ExaBotExaNot documentedNone published, user agent only
FirecrawlAgentFirecrawlYes, documentedNone published, user agent only
HenkBotUnidentifiedNot documentedNone published, user agent only
imageSpiderUnidentifiedNot documentedNone published, user agent only
Kangaroo BotKangaroo LLMNot documentedNone published, user agent only
laion-huggingface-processorLAIONNot documentedNone published, user agent only
LAIONDownloaderLarge-scale Artificial Intelligence Open NetworkNoNone published, user agent only
LCCUnidentifiedNot documentedNone published, user agent only
LightpandaUnidentifiedNot documentedNone published, user agent only
Mozilla-TabstackMozillaYes, documentedNone published, user agent only
MyCentralAIScraperBotUnidentifiedNot documentedNone published, user agent only
netEstate Imprint CrawlernetEstateNot documentedNone published, user agent only
omgiliomgilibotWebz.ioYes, documentedNone published, user agent only
PanguBotthe Chinese company HuaweiNot documentedNone published, user agent only
QueritBotQuerit-SearchBotQueritNot documentedNone published, user agent only
SemrushBot-OCOBSemrushYes, documentedNone published, user agent only
ShapBotParallelYes, documentedNone published, user agent only
SpiderUnidentifiedNot documentedNone published, user agent only
TavilyBotTavilyNot documentedNone published, user agent only
TerraCottaTerra CottaCeramic AIYes, documentedNone published, user agent only
VelenPublicWebCrawlerVelen CrawlerYes, documentedNone published, user agent only
WARDBotWEBSPARKNot documentedNone published, user agent only
Webzio-Extendedwebzio-extendedUnidentifiedNot documentedNone published, user agent only
YandexAdditionalYandexAdditionalBotYandexYes, documentedReverse DNS, forward confirmed

Other automated crawlers12

Bots that touch AI workflows without fitting the categories above.

Other automated crawlers: user-agent token, operator, robots.txt policy and how each one can be verified
CrawlerOperatorrobots.txtVerificationDetails
BuddyBotBuddyBotLearningNot documentedNone published, user agent only
EchoboxBotEchoboxNot documentedNone published, user agent only
facebookexternalhitMeta/FacebookNoOperator ASN check
Google-FirebaseGoogleNot documentedNone published, user agent only
OpenAIOpenAIYes, documentedPublished IP range file
Panscientpanscient.comPanscientYes, documentedNone published, user agent only
Poseidon Research CrawlerPoseidon ResearchNot documentedNone published, user agent only
QuillBotquillbot.comQuillbotNot documentedNone published, user agent only
SBIntuitionsBotSB IntuitionsYes, documentedNone published, user agent only
ThinkbotThinkbotNoNone published, user agent only
wpbotQuantumCloudNot documentedNone published, user agent only
YaKMeltwaterNot documentedNone published, user agent only

Check your own access log

Paste a hundred lines and every crawler claim in them is checked against what that operator publishes, the same way a single address is checked on a crawler's own page. Most logs come back with more unverifiable claims than false ones, which is its own finding: it is not that the traffic is proven fake, it is that nobody published a way to prove it real.

Questions this directory answers

Which AI crawlers ignore robots.txt?
9 of the 153 tokens listed here are documented as ignoring robots.txt: Bytespider, DeepSeekBot, facebookexternalhit, ISSCyberRiskCrawler, LAIONDownloader, Linguee Bot, Meta-ExternalFetcher, Perplexity-User, Thinkbot. A further 90 have no published policy either way, which is not the same as a promise to obey it. For those, a robots.txt rule is a request and an edge rule is enforcement.
How do you verify that an AI crawler is genuine?
By checking the source IP, never the user agent. Three things are checkable: a published IP range file, a forward-confirmed reverse DNS convention, or an operator ASN. Only 34 of the 153 tokens here publish any of them; for the rest there is nothing to check against.
Is the user-agent string enough to identify a crawler?
No. The user agent is a header the client chooses, so anything can send it. Scrapers copy the strings of well-known crawlers precisely because site owners allow those through. Treat the user agent as a claim and the IP as the evidence.
Does blocking a training crawler remove my content from the model?
No. Blocking stops future crawls; it does not remove anything already collected, and it has no effect on what a model already learned. It also does not change how the site ranks in ordinary search.
What is the difference between a training crawler and an AI search crawler?
A training crawler collects pages that end up in a model's training corpus. An AI search crawler builds the index an answer engine cites from. Blocking the first is the opt-out most publishers mean; blocking the second removes you from the answers.
Open data

Cite this dataset

Every row in this directory, 153 crawler tokens with their operator, user-agent string, robots.txt stance and verification method, as one file. Published under CC BY 4.0: use it anywhere, including commercially, as long as the credit line stays with it.

Dataset updated . The files are rebuilt from the same source as the table above, so the two never disagree.

IPScanner (2026). AI crawler directory. https://ipscanner.io/bot. CC BY 4.0.

Agentscan

Stop maintaining range files by hand

Every operator publishes its ranges in a different place, in a different format, and changes them without notice. Agentscan does the reverse-DNS and range matching for you and returns a verdict per request, so your edge can allow the real search crawlers and challenge everything wearing their names.