We asked four AI engines the same questions three times each. On the second ask, about half the sources they cited were gone.
When an AI assistant cites a source, it feels like a fact about the web: this is where the answer comes from. We tested how solid that fact is by asking the same questions again. It is not solid at all. We put the same questions to the four AI search engines that show their sources, three times each, and on the second ask about half the sources they had cited were simply gone, replaced by others. AI does not only disagree with rival engines. It disagrees with itself.
We took 20 everyday questions, the same kind people actually ask, and put each one to the four AI search engines that retrieve and show real web sources: Google’s Gemini (with search grounding), Microsoft’s Copilot (Bing), OpenAI’s ChatGPT search, and Perplexity. Then we did the one thing these tools are rarely tested on. We asked again. Each question went to each engine three separate times, and we traced every cited link to the real domain behind it, so we could compare exactly which sources came back on each identical run.
Averaged across every engine, only about half the sources survived to the next identical ask. The full source lists overlapped just 35% run to run. These are not different questions, or different engines, or different days. Same question, same engine, minutes apart, and half the evidence changed.
The clearest way to feel it is a single case. We asked Copilot "how to cook salmon" twice. The first time it cited eight sources. The second time it cited eight sources again, and not one of them was the same. A completely different set of websites, for the identical question, from the identical engine. That is the extreme rather than the average, but the average is bad enough: reask a question and a coin-flip’s worth of the sources you saw the first time are gone.
The instability is not evenly spread. Microsoft’s Copilot was the most volatile by a distance, changing about 60% of its sources on a re-ask. Perplexity was the most consistent, and it still changed 44%. Gemini and ChatGPT search landed in between. Put the other way round: the best-case chance that a source cited once gets cited again on the very next identical query was about 56%. Barely better than a coin toss, and for Copilot it was closer to two in five.
This is the quiet problem under every AI-visibility tool that hands you a single number. That number is built on a check run once. But a citation captured once has only about a 50% chance of being there the next time you look, so a one-off check is a photograph of a single dice roll, and a blended "score" is that photograph with a confident label on it. It can read impressively high or alarmingly low purely on which run it happened to catch. It is a guess with a decimal point.
The only way to say anything true about AI visibility is to measure it repeatedly, keep the actual citation as evidence, and watch the trend rather than a single reading. One capture tells you a source was reachable in one moment. A run of captures over time tells you whether you are genuinely, durably part of how an engine answers. That difference, evidence over time versus a one-shot score, is the whole reason Unsourced exists.
Treat any single AI-visibility reading, ours or anyone’s, as one sample, not a verdict. If a tool shows you one number and no history, it cannot tell the difference between a site AI reliably cites and one it happened to mention on the run it checked. Track each engine over time, keep the receipts, and look at how often you actually show up, not whether you appeared once. You can see a first reading for your own domain here, then watch whether it holds.
A citation looks like a fact, but on a grounded AI engine it is closer to a weather report: true for a moment, and quietly different the next time you check. Ask again and half of it moves. That does not make AI search useless, but it does make any single measurement of it fragile, and it makes one honest thing very clear. Being cited once is not being cited. Only the pattern over time is real.
No. We asked four AI search engines the same questions three times each, and on an identical re-ask about half the sources they had cited were gone. Overall only about 50% of cited sources repeated on a second run, and the two source lists overlapped just 35%.
Perplexity was the steadiest in our test, but steadiest still meant 44% of its sources changed on a re-run. Microsoft’s Copilot was the most volatile at about 60% churn. Gemini and ChatGPT search sat in between, at roughly 46% and 49%.
Grounded AI engines run a fresh web search each time and sample from a shifting set of results, and the model’s selection is non-deterministic. So the sources are a snapshot of one moment, not a fixed answer to the question.
That a single check or a one-off visibility score is close to meaningless. Being cited once has only about a 50% chance of repeating on the very next identical query. The only reliable read is repeated measurement over time, per engine, with the actual citation kept as evidence.
We asked 20 everyday questions to the four AI search engines that return real sources (Gemini with grounding, Copilot/Bing, ChatGPT search, Perplexity), three times each, resolved every cited link to its true domain, and measured how many sources carried over between identical runs. Snapshot, September 2026; every figure measured, not modelled.
If a citation barely survives a re-ask, a single number built on one has nothing under it.
Read it →The engines disagree with each other. This is each engine disagreeing with itself.
Read it →See which engines cite your domain right now, with the answer attached, then watch it over time.
Read it →Unsourced captures which AI assistants cite you, proves which crawlers really fetched your pages, and re-checks after you act — evidence, not a score.
© Unsourced, the evidence layer for AI search. Our data and findings are free to quote and cite — please attribute to Unsourced and link to unsourced.app.