An AI agent just requested a page and the log line says GPTBot/1.2. That string is not
evidence. A User-Agent is written by the client, so anyone can send it, and the only reliable
way to verify an AI agent is to check the address it connected from. Three checks do the real
work: match the source IP against the operator's published range file, run a forward-confirmed
reverse DNS lookup, or confirm the IP is announced by the operator's own ASN.
Which of the three applies depends entirely on the operator, and that is the annoying part. OpenAI publishes JSON range files per bot. Anthropic publishes one prefix file covering ClaudeBot, Claude-User and Claude-SearchBot together. Apple supports both reverse DNS and a range file. Meta publishes neither and expects you to check AS32934. Plenty of crawlers publish nothing at all, at which point the name in the header is decoration and you decide on origin and behaviour instead.
The claim is free, the check is not
Think about what an allowlist rule like "if User-Agent contains GPTBot, skip the bot rules" actually says. It says: any client willing to type nine characters gets the treatment you reserved for a crawler you trust. Scrapers read your robots.txt, notice which agents you welcome, and dress accordingly. The impersonation is not clever, it is just cheap, and it works against any site that never looks past the header.
The check costs more. You need the operator's current ranges, a DNS resolver, or routing data, per operator, refreshed as they change. Researchers have already reported traffic claiming to be PerplexityBot from addresses outside Perplexity's published ranges, which is exactly the failure mode: the name was borrowed, the network was not.
Check which network an agent request really came from
Every operator documents this differently
There is no shared standard, so verification is a per-operator lookup. A sample from the AI crawler directory, which tracks 150 agents:
| Agent | Operator | What they publish | How you verify |
|---|---|---|---|
| GPTBot | OpenAI | openai.com/gptbot.json | IP range match, no reverse DNS convention |
| ClaudeBot | Anthropic | claude.com/crawling/bots.json | IP prefix match, one file covers all Anthropic bots |
| PerplexityBot | Perplexity | perplexitybot.json | IP range match, spoofing already reported in the wild |
| Applebot | Apple | Range file plus rDNS | Reverse DNS under applebot.apple.com, forward confirmed |
| Amazonbot | Amazon | Reverse DNS convention | Forward-confirmed reverse DNS, no range file |
| Meta-ExternalAgent | Meta | Nothing per bot | Confirm the IP is announced by AS32934 |
| Bytespider | ByteDance | Nothing | Not verifiable, and it ignores robots.txt |
| CCBot | Common Crawl | Nothing usable | Not verifiable, runs from rented cloud capacity |
| Google-Extended | Robots token only | Nothing to verify, it never appears in a User-Agent |
Google-Extended is worth a second look because it breaks the mental model. It is a robots.txt directive for opting out of Gemini training, not a crawler. If a request arrives claiming to be Google-Extended, it is a forgery by definition. There is no real version of it.
Of the 150 agents in the directory, 33 expose a usable check through published ranges, reverse DNS or an ASN. Nine are documented as ignoring robots.txt outright. The rest sit in between, publishing a name and nothing to back it up.
What impersonation looks like in your logs
Two requests, same header, minutes apart:
198.51.100.7 "Mozilla/5.0 ...; compatible; GPTBot/1.2; +https://openai.com/gptbot"
203.0.113.44 "Mozilla/5.0 ...; compatible; GPTBot/1.2; +https://openai.com/gptbot"
Nothing in the line separates them. Resolve the addresses and the difference is immediate: one falls inside OpenAI's published egress ranges, the other is announced by a hosting provider that rents servers by the hour. The second one also requested 400 product pages in ninety seconds and never loaded a stylesheet, which no crawler that respects its own published crawl budget does.
The pattern repeats across operators. A fake ClaudeBot from a VPS range. A fake Applebot with a PTR record that resolves to something ending in a hosting company's domain, or with no PTR at all. A fake Amazonbot whose PTR looks plausible until the forward lookup returns a different address, which is why the forward confirmation step exists. In every case the header matched perfectly and the network did not.
Verified, unverified and hostile are three different answers
Once you stop treating the header as proof, requests fall into groups that deserve different handling, and a yes/no bot flag cannot express that.
A verified AI crawler is a policy decision: allow it, meter it, or block it by name because you do not want your content in that corpus. A request claiming a name that fails the IP check is not a policy decision, it is a forgery, and it should never inherit the access the real agent gets. An unverifiable crawler with no published anything needs rate limits rather than a verdict. And a headless browser running from a datacenter range with no bot claim at all is the case your allowlist was never looking at.
That is why Agentscan returns a class rather than a boolean. Every request comes
back as human, known_bot, ai_agent or malicious_automation, with a confidence value and
the signals that produced it.
The signals that survive a fake User-Agent
IP origin answers most of the question, but not all of it. An agent can run from a residential address, and a scraper can rent a clean IP with no history. Three more signals hold up when the header lies:
JA4 TLS fingerprint. The client reveals how it speaks TLS before it sends a single header: cipher order, extensions, ALPN, version. A Python script announcing itself as Chrome 124 produces a handshake that no Chrome build has ever produced. The header is a claim, the handshake is behaviour.
Headless tells. The webdriver flag, missing browser surfaces, and automation markers left
by Playwright, Puppeteer or a raw HeadlessChrome build. On its own each one is weak. Combined
with a datacenter origin, the picture is not ambiguous.
Header consistency. Real browsers send Accept, Accept-Language and Accept-Encoding as a set, in a stable order, matching the version they claim. Hand-built request objects usually send a subset, and the gaps are consistent enough to be a signal.
None of these is decisive alone, which is the point. Fusing them is what lets you separate a verified crawler from a competent impersonator, and both from an ordinary visitor. The same stacking logic applies to older crawlers too, covered in how to verify Googlebot.
A policy that survives contact with real traffic
Verification is only useful if something happens afterwards. A workable default:
- Verify the claim against the operator's method before any allowlist rule fires. Range file, reverse DNS or ASN, whichever that operator publishes.
- Failed verification is not a soft signal. A request claiming to be ClaudeBot from outside Anthropic's prefixes is impersonation. Treat it as you would any unidentified scraper, not as a maybe.
- Split the allow decision by purpose. Training crawlers, AI search crawlers and assistant fetchers deserve separate rules. Blocking a search crawler removes you from that engine's cited answers, which is a traffic decision, not a security one.
- Rate limit the unverifiable. For Bytespider, CCBot and everything else with no published check, the name tells you nothing. Request rate, path pattern and origin tell you plenty.
- Log the class, not just the block. Knowing that AI agent traffic tripled last month is worth more than a counter of blocked requests, especially when you are deciding what to monetise.
Agentscan runs this from one POST to /v1/agentscan/check: reverse-DNS-verified allowlist,
published range matching, IP origin, headless flags and JA4 in a single call, cached so repeat
lookups come back in under 50ms. Fast enough to sit in middleware rather than in a nightly log
review.
Bottom line
You cannot verify an AI agent from what it calls itself. Check the source IP against whatever the operator actually publishes, treat a failed check as a forgery, and use JA4, headless tells and IP origin to classify everything the allowlist cannot vouch for. The question was never "is this a bot". It is which of the four kinds of traffic just arrived, and what you want to do about each one.
FAQ