We checked every request that reached our network wearing an AI crawler's name for 30 days, against each operator's published IP ranges and forward-confirmed reverse DNS. More than one in five failed. The failures were not random noise: the same IP addresses turned up claiming to be five different AI companies at once. A user-agent is a claim, and a large share of the ones we saw were lies.
For 30 days we logged every request that reached our network carrying an AI crawler's user-agent — GPTBot, ChatGPT, OAI-SearchBot, PerplexityBot, Google-Extended and the rest — and checked each one against two things the sender cannot fake: the IP ranges the operator publishes for its crawler, and forward-confirmed reverse DNS to the operator's own domain. 21.8% of them failed both. More than one in five requests wearing an AI crawler's name did not come from that AI company.
The failures were not random misconfigurations. The giveaway that this is deliberate is that the same IP addresses showed up under multiple operator names. Five separate addresses each appeared as ChatGPT and GPTBot and OAI-SearchBot and PerplexityBot and Google-Extended — all five, from one address. A single Google Cloud VM accounted for 323 of these hits on its own, rotating through five identities.
No real crawler is OpenAI and Perplexity and Google at the same time. This is one scraping operation rotating AI-bot names to get past the filters that wave "good AI bots" through. The AI-crawler label has become a skeleton key, and people are using it.
A user-agent string is written by whoever sends the request, so on its own it proves nothing. What a spoofer cannot reproduce is the operator's network footprint. We check every claim two independent ways:
This is why the check is honest in both directions. Real crawlers pass it: genuine GPTBot and ClaudeBot requests from the operators' own ranges verified cleanly and were never flagged. Only the requests that could prove neither were counted as impostors. We do not accuse a bot we simply cannot verify — the 21.8% is requests that contradict their own claim, not requests we were merely unsure about.
The common advice is to allow the "good" AI crawlers so your content feeds the assistants. That advice quietly assumes the name in the header is true. If you allow-list GPTBot by its user-agent, you are not allow-listing OpenAI — you are allow-listing anyone who types "GPTBot," including the scraper that also claims to be Perplexity and Google. Any rule keyed on the header alone can be walked straight through.
This is our own monitored traffic, one network, over a single 30-day window. It is not a census of the whole web, and the rate will differ from site to site. But the shape of it — coordinated spoofing across AI-bot names, from shared infrastructure — is not something we would expect to be unique to us, and in our data it is climbing: it was closer to one in six a week earlier.
You can check a single request with our free AI crawler verifier, and for the wider picture of who is really crawling real sites, see the 2026 AI Crawler Report. For how spoofing works and the three verdicts a crawler can get, see how AI crawler spoofing works.
"AI crawler" is now a name worth stealing, and more than one in five of the ones we logged were stolen. The user-agent was never proof; it is just the easiest part to fake. The only reliable question is whether a request can back its claim with an address the operator actually owns. Ours can. A header cannot.
It sent a request carrying a real crawler's name — GPTBot, PerplexityBot, ClaudeBot and so on — from an address that is not that operator's. We judge this against the IP ranges the operators publish and forward-confirmed reverse DNS, not the header. A request that matches neither is wearing the name, not using it.
Because the crawler name lives in the User-Agent header, which the sender writes. Nothing stops one machine from sending a request as GPTBot, then another as PerplexityBot, then another as Google-Extended. In our data, five separate addresses did exactly that across all five identities — one scraping operation cycling AI-bot names to slip past filters that wave "good bots" through.
No — the real ones can be worth allowing, and they verify cleanly. The point is that allow-listing by user-agent alone allows whoever types the name. Allow the crawlers you want by verified identity (published range plus reverse DNS), not by the header they send.
Match the source IP against the operator's published range and run a forward-confirmed reverse-DNS lookup to their domain. OpenAI, Google, Perplexity and now Anthropic all publish their ranges. You can check a single request with our free crawler verifier, and Unsourced keeps the running record across engines.
Check whether a bot name or IP is really who it claims.
Read it →Why a user-agent proves nothing, and the three verdicts a crawler can get.
Read it →Who is really crawling real sites, verified against published ranges.
Read it →Unsourced captures which AI assistants cite you, proves which crawlers really fetched your pages, and re-checks after you act — evidence, not a score.
© Unsourced, the evidence layer for AI search. Our data and findings are free to quote and cite — please attribute to Unsourced and link to unsourced.app.