Every AI-visibility tool hands you a number out of 100 and calls it your visibility. Before you trust it, or pay for it, it is worth seeing exactly how that number gets made. Once you see the machinery, you cannot unsee it.
When someone asks ChatGPT or Perplexity a question, the model does not look up one correct answer. It generates a fresh answer on the spot, and many engines run a live web search first and fold in whatever they find that second. Two things follow, and they matter for everything below. The answer is not fixed: ask again and it can change. And each engine runs its own search over its own index, so different engines read different pages entirely.
A scoring tool cannot ask every question, so it picks a sample. It chooses prompts it thinks your buyers use, asks a few AI models each one, checks whether your name showed up, weights the results with a formula it designed, and blends everything into a single number. That number is the score.
Read that back slowly, because the whole thing rests on it. A visibility score is a sample of questions the tool chose, fired at a moving target, checked for your name, and averaged by a formula you cannot see. It is an estimate of an estimate. And it blends together engines that, as it turns out, barely agree with each other.
We measured this. We asked the three AI engines that actually show their sources, Google's Gemini, Perplexity, and ChatGPT with search, the same 156 questions, and recorded the exact domains each one cited. Any two engines cited the same source only about 8% of the time. 91% of what an engine cited, no other engine named. Here is one real question from that study. Click through the three engines and watch what each one reaches for.
One real question from the study: “What’s the best managed Postgres provider for a small SaaS?”
Same question. Three engines. Not one source in common. A single “visibility score” blends all three into one number and hides exactly this.
Same question. Three engines. Not one source in common. That is not a trick query, it is the normal case. So when a tool blends those three engines into one number, it is averaging three different internets. The number can go up while you vanish from the one engine your buyers actually use.
A score is not useless. As a rough gut-check, "are we roughly present or roughly absent in AI answers", it is fine. The trouble starts when you ask it to carry weight it was never built for. A visibility score cannot tell you whether the AI read your page or just recognised your name. It cannot tell you which engine, on which day, cited you. It cannot tell you which competitor was named instead. And it cannot be audited, because you cannot open it and check the working. The moment being cited starts to matter, a licensing conversation, a rights claim, a client renewal, or just knowing whether a change you made worked, a number that "went up" does not survive the first hard question.
The alternative to a score is not a better score. It is a receipt. Instead of blending everything into one number, you keep the evidence:
None of that is a number you take on faith. It is a record you can open, hand to a partner, or attach to a claim. If we cannot prove it, we do not put a number on it.
All figures and examples are accurate as of publication. AI answers are generated live and shift over time, so the exact sources shown above will change; the low overlap between engines is the stable finding.
No, and that is not the point. A score is a real estimate from a real sample. The issue is what it is asked to do. As a rough sense of whether you are present in AI answers, a score is fine. As proof that a specific engine cited your specific page, it cannot do the job, because it is a blended average you cannot audit, and the engines it blends agree only about 8% of the time.
Because there is no single "AI." We measured 156 questions across the three engines that show their sources and they cited the same domain only about 8% of the time. You can be cited by one engine and invisible on another. Averaging them into one number hides the exact gap you would want to fix.
A rough, directional sense of whether you show up in AI answers at all. That is genuinely useful as a starting point. What it cannot do is prove which engine cited you, on what date, from which page, or whether a competitor was named instead. Those need evidence, not an estimate.
Keep the evidence, per engine: the exact answer, the date, the source the engine pulled from, and confirmation that a verified crawler actually fetched your page. That record can be opened, checked, and handed to someone else. A score cannot. You can see the difference on your own domain with the free AI Visibility Check.
The full measured data behind why one number cannot be right.
Read it →See how many different sites five AI models name for your topics. No login.
Read it →Check any AI crawler identity against its operator published ranges.
Read it →Unsourced captures which AI assistants cite you, proves which crawlers really fetched your pages, and re-checks after you act — evidence, not a score.
© Unsourced, the evidence layer for AI search. Our data and findings are free to quote and cite — please attribute to Unsourced and link to unsourced.app.