Free tool
robots.txt generator for AI crawlers
Pick the crawlers you want to keep out and this writes the robots.txt rules for them, aliases included. It covers all 153 AI crawler tokens in the directory and shows, next to each checkbox, whether that crawler can be verified at all. That is the part other generators leave out, and it is the part that decides whether a rule does anything: only 34 of these operators publish a way to prove a request is really theirs, so for the other 119 the rule stops whoever chooses to be stopped.
What robots.txt can and cannot do
robots.txt is a published request. A crawler reads it, and a well-behaved one does what it says. Nothing about the file enforces anything: there is no authentication in it, no way to tell a real GPTBot from a scraper that typed the same string into a header, and no penalty for ignoring it. 9 of the crawlers in this directory are documented as ignoring the file outright, and 90 have never said either way.
That is still a reason to write one rather than a reason not to. A robots.txt is the record of what you asked for, and every operator that has committed to honouring it will. It is also the thing a CDN rule, a licensing conversation or a legal argument points back at later. What it is not is a wall, and a generator that hands you one without saying so has done you a disservice.
One practical note before you paste. Blocking AI crawlers has nothing to do with how the site ranks in ordinary search: Googlebot and Bingbot are search crawlers, they are not in this list, and nothing here touches them. The rule that does cost you something is blocking an AI search crawler, because that removes the site from the answers those engines write. Opting out of training and opting out of being cited are two different decisions, and the presets below keep them apart.
Build the file
Crawler list last updated . Nothing you tick here leaves your browser.
robots.txt and AI crawlers, answered
- How do I block AI crawlers in robots.txt?
- Add one rule group per crawler token: a User-agent line naming the bot, then Disallow: / underneath it. There is no wildcard that means "AI crawlers", so a real opt-out means naming all 153 of the tokens in use, aliases included. That is what this generator writes.
- Does blocking AI crawlers hurt my SEO?
- No, as long as you block the right ones. Googlebot, Bingbot and the other ordinary search crawlers are not AI crawlers and are not in this list, so nothing here touches how the site ranks. The rule that does cost you traffic is blocking an AI search crawler: that removes the site from the answers those engines write, which is a different decision from opting out of model training.
- What is the difference between blocking training and blocking AI search?
- Training crawlers collect pages that end up in a model's training corpus; AI search crawlers build the index an answer engine cites from. Blocking the first is the opt-out most publishers mean, and covers 61 tokens here. Blocking the second removes you from the answers. The default preset on this page blocks training and dataset collection and leaves answer engines alone.
- Will AI crawlers actually obey robots.txt?
- Some will and most are unproven. 9 of the 153 tokens listed here are documented as ignoring robots.txt outright, and 90 have published no policy either way, which is not the same as a promise to obey it. robots.txt is a request; for anything you need enforced, the check has to happen at the request itself.
- Why does the generator show a verification status for each crawler?
- Because it decides whether the rule can be enforced. Only 34 of the 153 tokens publish an IP range file, a reverse-DNS convention or an operator ASN you can check against; the remaining 119 can be typed into any user-agent header by anyone. Blocking those is still worth doing as a statement of intent, but it stops whoever chooses to be stopped.
- Do I need to list the aliases as well as the main name?
- Yes. robots.txt matches on the user-agent token, so a rule naming ClaudeBot does nothing to a request identifying as Claude-Web. Every rule this generator writes covers the crawler's aliases alongside its current name, which is why the output is longer than the number of crawlers you ticked.
- Where does the robots.txt file go?
- At the root of the domain, as https://example.com/robots.txt, served as plain text. Rules are per host and per protocol, so a subdomain needs its own file. If you already have one, paste these groups into it rather than replacing it: robots.txt has no include mechanism and the last file wins.
Look a crawler up first
Every entry in the directory says who runs the crawler, what the token looks like in a log, whether the operator has committed to robots.txt, and how a request claiming to be it can be checked.
When the request has to be checked, not asked
A robots.txt rule is read by the crawler. Agentscan reads the request instead: it matches the source against published crawler ranges, forward-confirmed reverse DNS and hosting data, and returns a verdict your edge can act on. That is what turns the file you just generated into something enforced.