Unsourced
Study

Same question. Three AI engines. 8% overlap.

Each one cites a different internet.

Ask ChatGPT, Perplexity and Google's AI the same question and you would expect a lot of overlap in what they cite. There isn't. We measured 156 questions across the three engines that actually show their sources, and on average any two of them agreed on about 8% of the domains they cited. Most of the time they shared nothing at all. Each engine is, in effect, its own internet.

Published 11 August 2026 · Unsourced Research · All research

The one-question version

Ask the three AI search engines that disclose their sources, Google's Gemini with grounding, Perplexity, and ChatGPT with search, a single question: "What's the best managed Postgres provider for a small SaaS?"

Gemini reaches for Neon, Supabase and Horizon. Perplexity reaches for Railway, Northflank and Seenode. ChatGPT reaches for Bytebase, Stackblaze and Techplained. Three engines, three completely separate lists, not one domain in common. That is not a trick question. It is the normal case.

What we measured

We took 156 real questions and put each one to the three engines above, then recorded the exact set of web domains each engine cited in its answer. These are not competitors we inferred from the prose. They are the sources the engines themselves returned in their grounding metadata, the receipts behind each answer. For every question we compared the three source sets and measured how much they overlapped. Our methodology sets out how we grade a citation from a passing mention to a proven fetch.

The corpus combines two things: a purpose-built batch of topically diverse questions we ran for this study (home services, health, finance, consumer tech, B2B software, travel, legal, automotive and more), together with a comparable set of questions from our own historical citation scans. We report both separately below, because the honest test of a finding is whether a sample we built on purpose agrees with data we did not.

What we found

The purpose-built batch, on its own, came in at 8.5% overlap, 90% unique sources and 37% zero-overlap questions. That is within a rounding error of the historical data. A deliberately diverse sample and an incidental one told the same story, which is the result worth trusting.

It holds on serious questions, not just product picks

It would be easy to assume this only happens on noisy "best tool" queries. It does not. Ask "Is intermittent fasting effective for weight loss?" and Gemini grounds on Harvard and the NIH, Perplexity grounds on Cochrane and the Mayo Clinic, and ChatGPT grounds on PubMed and Frontiers. Every source is credible. Not one is shared. The engines are not disagreeing on quality. They are drawing from different corners of the web entirely.

The same held for "How do I choose a business bank account for a new limited company?", "Which continuous glucose monitor is best for a non-diabetic?", and "What's the best accounting software for a UK sole trader?" In each, all three engines answered, and all three cited a different set of sources.

Why the engines diverge

Because there is no "the AI." Each engine runs its own retrieval stack: its own crawl and index, its own live-search backend, its own ranking of which pages to pull into an answer. Gemini leans on Google's index. Perplexity runs its own search and blends in sources like Reddit. ChatGPT's search tool has its own retrieval path again. Same question, three different machines deciding what to read. Low overlap is not a fault in any of them. It is the direct consequence of three independent systems each answering from their own view of the web.

What this means if you are trying to get cited

Two things follow, and both cut against how "AI visibility" is usually sold.

First, a single blended "AI visibility score" hides the fact that matters most: the engines disagree. If you are cited by Perplexity and invisible on Gemini, a number that averages the two tells you nothing you can act on. The useful view is per engine, every time.

Second, chasing "more prompts" or "the right prompts" is a treadmill. When the engines do not even cite the same sources for the same question, there is no universal set of prompts that reveals your true standing. What you can act on is narrower and more honest: for a specific engine, on a specific question, were you cited, who was cited instead, and did a verified crawler from that engine actually fetch your page. That is measurable. A blended score is not.

What this does not prove

We want to be precise about the limits, because a study that overclaims is exactly the thing we built Unsourced to push back on.

The takeaway

There is no single "AI" citing you or ignoring you. There are several, and they barely agree. "Am I cited?" answered on one engine tells you almost nothing about the others. If you want the real picture, you have to look at each engine on its own, with the receipts, rather than average them into a number that quietly hides the disagreement. That is the whole idea behind Unsourced: evidence, per engine, not a score.

Common questions

Do different AI engines cite the same sources for the same question?

Mostly not. Across 156 questions put to the three engines that disclose their sources, any two engines cited the same domain only about 8% of the time, and 91% of all cited sources appeared in just one engine. On 38% of questions the engines shared no source at all.

Why don't ChatGPT, Perplexity and Gemini cite the same sources?

Because there is no single "AI." Each engine runs its own retrieval: its own crawl and index, its own live-search backend, its own ranking of which pages to pull into an answer. Gemini leans on Google's index, Perplexity runs its own search and blends in sites like Reddit, and ChatGPT's search tool has its own path again. Three independent systems reading different pages produce different citations.

Does this mean a single AI visibility score is misleading?

Yes. If the engines agree only about 8% of the time, averaging them into one number hides the fact that matters most: you can be cited by one engine and invisible on another. The only honest view is per engine. A blended score smooths away the exact gaps you would want to act on.

How did you measure this?

We sent the same questions to Google Gemini with grounding, Perplexity, and ChatGPT with search, then captured the exact set of domains each engine cited from its own grounding metadata, the sources the engine itself returned rather than anything we inferred. For each question we compared the three source sets and measured the overlap. The corpus combines a purpose-built batch of topically diverse questions built for this study and our historical citation scans, which independently agreed.

Why only three engines, not ChatGPT, Claude, Grok and the rest?

Only some engines return a citable list of the sources they read. Gemini with grounding, Perplexity, and ChatGPT with search do. The others, including Claude, Grok, Llama and a plain non-search GPT, answer from training data and return no source list, so their citations cannot be measured this way. This study is about the engines that show their work.

Related

A score is a guess. Ask for the receipt.

If the engines disagree this much, one blended number cannot be telling you the truth.

Read it →

How to check if AI cites your website

The per-engine way to see where you actually stand across ChatGPT, Perplexity, Gemini and Grok.

Read it →

See the receipts in the live demo

Real captured answers and the sources each engine named. No login.

Read it →

See the evidence for your own site

Unsourced captures which AI assistants cite you, proves which crawlers really fetched your pages, and re-checks after you act — evidence, not a score.

© Unsourced, the evidence layer for AI search. Our data and findings are free to quote and cite — please attribute to Unsourced and link to unsourced.app.