Unsourced
Original Research · Field Report · 2026

How much of your ‘AI crawler’ traffic can you actually verify?

We analysed 14,227 real AI-crawler visits, checking each one against the operator's true network identity: its own live-published IP ranges and forward-confirmed reverse DNS. The result is a clear picture of which AI crawlers can prove who they are, and why an IP outside a published range is not automatically fake.

14,227 crawler visits · 18 AI crawlers · ~2,015 unique IPs · 10 Jun–6 Aug 2026 · checked against operators' live-published IP ranges + forward-confirmed reverse DNS · every figure measured, none simulated

This is the 2026 edition. It expands our first crawler report with a refined method and a larger dataset.

56%
verified: the IP is inside the operator's live-published range, or forward-confirmed by reverse DNS
34%
outside a published range: claimed an operator that publishes ranges, from an IP not on its list. Not proof of fakery (see below)
10%
unverifiable: the operator publishes nothing to check against, so we flag but never accuse
Verified 56% Outside a published range 34% Unverifiable 10%

A verified visit proved its identity by network, not by its user-agent label. Outside a published range means the operator publishes a list and this IP wasn't on it. A flag, not a verdict of fraud. Unverifiable means the operator publishes nothing to check against; we never count it against anyone. Measured by visit; by distinct IP the split is 59% / 27% / 14%.

The real divide in 2026: who lets you verify them

The useful split is not real-versus-fake. It is which operators publish the means to verify their crawlers at all.

The news this year: Anthropic changed sides. For a long time its guidance said it did not publish IP ranges. In 2026 it began serving a public feed for ClaudeBot, Claude-User and Claude-SearchBot. The moment it did, ClaudeBot traffic in our data went from effectively unverifiable to 92% verified. Publishing the means to check is a choice, and the list of operators making it is growing.

The crawlers we saw, and how much of each verified

The busiest crawlers we saw, by visit count, with each one's verified / outside-range / unverifiable split and whether its operator publishes ranges. The full 18-crawler table is in the downloadable dataset.

AI crawlerVisitsVerifiedOutside rangeUnverifiablePublishes ranges
ChatGPT-User (OpenAI, on-demand)7,88853%47%0%Yes
ClaudeBot (Anthropic)1,96292%8%0%Yes (new in 2026)
Bytespider (ByteDance)1,0620%13%87%No
GPTBot (OpenAI, training)97182%18%0%Yes
OAI-SearchBot (OpenAI, search)94669%31%0%Yes
PerplexityBot (Perplexity)47766%34%0%Yes
Amazonbot (Amazon)35762%2%36%Yes (via rDNS)
CCBot (Common Crawl)3170%4%96%No
↓ Download report (PDF) ↓ Download dataset (CSV) ↓ Download dataset (JSON)

Each row is one AI crawler. Columns: operating company, type (search, trainer or scraper), total visits, distinct IP addresses, percent verified, percent outside the operator's published range, percent unverifiable, and whether the operator publishes IP ranges you can check against. Aggregate only, with no per-request or per-site data. Free to cite with attribution to Unsourced.

Why ‘outside the range’ is a flag, not a verdict

This is the distinction most tools miss, and it is the reason to read the network instead of the label. When a crawler's IP sits outside its operator's published list, that is worth flagging. On its own, it is not proof of fraud, and treating it as such is how you end up accusing a real operator of impersonating itself.

So we report “outside the published range” as exactly that, and we reserve “fake” for the traffic that earns it. The figure worth acting on is the verified one. Drawing that line cleanly, rather than inflating a scary headline, is the whole point of measuring identity properly.

What this actually means for your logs

If a line in your server log says “GPTBot” or “ClaudeBot,” that looks like proof an AI read your page. It isn't. It is a label any script can type. Here is the honest hierarchy of proof:

The takeaway: verify the visitor, then trust the log, never the other way round. Turning a user-agent claim into a checked verdict, on your own site, with the evidence attached, is exactly what Unsourced does.

How we measured this. Every AI-crawler visit Unsourced recorded between 10 June and 6 August 2026 (14,227 visits, roughly 2,015 unique IPs, 18 named crawlers) was resolved to a verdict using two independent proofs: whether the IP falls inside the operator's live-published ranges, and forward-confirmed reverse DNS (the IP's reverse record must resolve to the operator's domain and that host must resolve back to the same IP. A PTR record alone, being attacker-controllable, is not trusted). Verified = at least one proof matched. Outside a published range = the operator publishes a list and the IP was not on it, reported as-is, never asserted as fraud. Unverifiable = the operator publishes no ranges and offers no confirmable reverse DNS, so no decisive check exists; we never label these fakes. Published-range coverage as of this report: OpenAI, Anthropic, Apple, Perplexity and Google. This is aggregated by operator, not an internet-wide census. We report the exact figures and let them stand.

See who's really crawling your site

Unsourced catches AI crawlers server-side and proves each one's identity as verified, outside-range or unverifiable, then hands you the evidence as a downloadable log you can filter, keep, or attach as proof.

© Unsourced. The evidence layer for AI search. Figures aggregated by operator. Full dataset linked above. Data and findings free to quote and cite — please attribute to Unsourced and link to unsourced.app.