Free tool
Pick a preset or tick crawlers. All 153 tokens, aliases included, built in your browser.
# Training crawlersUser-agent: AI2BotUser-agent: Ai2Bot-DolmaUser-agent: anthropic-aiUser-agent: Applebot-Extended# 19 moreDisallow: /
# AI search crawlersUser-agent: AddSearchBotUser-agent: AIWebIndexUser-agent: AIWebIndex-AgentUser-agent: amazon-kendra# 26 moreAllow: /
# All 153 crawlersUser-agent: AddSearchBotUser-agent: AgentTimesUser-agent: AI2BotUser-agent: AI2Bot-DeepResearchEval# 166 moreDisallow: /
Crawler list updated . Nothing you tick leaves your browser.
Before you paste
Only 34 of 153 crawlers can be verified. Any other name can be typed by anyone.
9 are documented as ignoring robots.txt. 90 never said either way.
Googlebot and Bingbot are not on this list. Blocking AI search does remove you from AI answers.
FAQ
Add one rule group per crawler token: a User-agent line naming the bot, then Disallow: / underneath it. There is no wildcard that means "AI crawlers", so a real opt-out means naming all 153 of the tokens in use, aliases included. That is what this generator writes.
No, as long as you block the right ones. Googlebot, Bingbot and the other ordinary search crawlers are not AI crawlers and are not in this list, so nothing here touches how the site ranks. The rule that does cost you traffic is blocking an AI search crawler: that removes the site from the answers those engines write, which is a different decision from opting out of model training.
Training crawlers collect pages that end up in a model's training corpus; AI search crawlers build the index an answer engine cites from. Blocking the first is the opt-out most publishers mean, and covers 61 tokens here. Blocking the second removes you from the answers. The default preset on this page blocks training and dataset collection and leaves answer engines alone.
Some will and most are unproven. 9 of the 153 tokens listed here are documented as ignoring robots.txt outright, and 90 have published no policy either way, which is not the same as a promise to obey it. robots.txt is a request; for anything you need enforced, the check has to happen at the request itself.
Because it decides whether the rule can be enforced. Only 34 of the 153 tokens publish an IP range file, a reverse-DNS convention or an operator ASN you can check against; the remaining 119 can be typed into any user-agent header by anyone. Blocking those is still worth doing as a statement of intent, but it stops whoever chooses to be stopped.
Yes. robots.txt matches on the user-agent token, so a rule naming ClaudeBot does nothing to a request identifying as Claude-Web. Every rule this generator writes covers the crawler's aliases alongside its current name, which is why the output is longer than the number of crawlers you ticked.
At the root of the domain, as https://example.com/robots.txt, served as plain text. Rules are per host and per protocol, so a subdomain needs its own file. If you already have one, paste these groups into it rather than replacing it: robots.txt has no include mechanism and the last file wins.