AI Visibility Score Reliability and What Makes Them Misleading
These single-point measurements can't distinguish brand progress from random AI variance.

AI visibility scores appear in board decks now, printed to a decimal point, sitting next to other core growth metrics like they belong there. That precision is misleading. Most of these numbers come from probabilistic systems tested with a small, vendor-chosen set of synthetic prompts, snapshotted at a single moment. They aren't a stable measurement of brand presence. They're one sample from a process that produces a different answer nearly every time you run it.
The generative engines behind these scores don't store a ranked list of results the way a search index does. Each answer gets built fresh, token by token, from a probability distribution over what should come next. Run the same prompt twice on the same day and you can get different brands named, different sources cited, different phrasing entirely. A paper out of the University of St. Gallen, posted to arXiv in April 2026, makes the case that "stability" needs to sit alongside "visibility" as a required metric. The paper's argument: a brand's generative engine optimization performance is a distribution across repeated runs, not a single-point outcome. Score it once and you've measured a coin flip, not a trend. A move from 8% to 11% between two reporting periods, under that framing, can't be told apart from noise. That gap sits inside the volatility that's baked into how these systems generate text in the first place.
None of this argues against measuring AI visibility at all. It argues for knowing exactly what the number in front of you describes before a marketing team builds a quarter's strategy around it.
How the engines themselves introduce variance that no dashboard can fully account for
Run the same query against Gemini repeatedly and the source list barely holds together: an arXiv study covered by Forbes in August 2026 found roughly 30% source overlap on Gemini across repeated identical queries, 33% to 40% on SearchGPT, and about 50% on Perplexity. Flip that around and it means more than half the cited sources on Gemini change from run to run, even when nothing about the query or the brand changed at all.
Purchase-intent prompts, the queries brands actually care about because someone's close to buying, turn out to be the least stable of the bunch. Analysis of 14,000 API calls found only about 4 in 10 brands appeared consistently across two runs of the identical prompt. That's the highest-value query type for most companies, and it's also the shakiest ground to build a KPI on.
Model updates make things worse in a different way. When a major model update shifts how a platform decides what to cite, marketing teams watching their dashboards see sharp drops in reported visibility. Nothing about the brand changed. The platform changed underneath it.
Interface mode matters too, and this one rarely gets discussed. Research examining ChatGPT's fast Instant mode versus its slower Thinking mode found that cited sources overlapped only a fraction of the time for the exact same prompt. Most tracking tools query one mode and report it as "ChatGPT," full stop.
Then there's persona. One index covering 126 software categories found a single brand's answer share swing from 34% to 90% depending entirely on whether the user described themselves as a freelancer or an enterprise buyer. Same brand, same category, wildly different outcome, based on a detail that has nothing to do with the brand's actual market position.
And the prompt being tracked often isn't the prompt that runs. AI engines decompose a user's question into multiple sub-queries that execute in parallel, a process called query fan-out. The "stimulus" a monitoring tool thinks it's measuring simplifies what the engine actually did behind the scenes.
Add it up and the pattern is clear: this variance isn't a rough edge that better engineering fixes next quarter. It's native to how these systems work.
The hidden denominator that makes any single percentage misleading
Search visibility used to rest on solid ground. A keyword set was finite, defined, and measurable, with a search volume number attached to it. You knew the denominator.
AI prompts don't work that way. Users type conversational, open-ended questions with effectively unlimited variation in wording, framing, and intent. There's no finite keyword list to anchor against.
So what does a visibility score actually measure? Brand presence inside a synthetic test set, usually dozens to a few hundred prompt permutations, chosen by whichever vendor built the dashboard. That's a small, curated slice of a space that has no real edges.
The IAB released a measurement framework in August 2026 that draws a line here: programs running fewer than 50 queries don't even qualify as directional. That's a floor, and plenty of commercial dashboards on the market today don't clear it.
The problem of a shrinking denominator gets worse because it's invisible. A score is computed against one vendor's prompt library, not against the full universe of things users actually ask. Two different tracking tools can report two different visibility numbers for the same brand on the same day, and both numbers can be technically accurate, because they're measuring against two different, undisclosed denominators. What looks like a market share figure is really a performance-within-a-test figure. The test itself rarely gets shown to the person reading the score.
What "mention," "citation," and "recommendation" measure, and why conflating them inflates scores
Most dashboards roll every kind of brand appearance into one aggregate score. That's a category error, because a mention, a citation, and a recommendation are not the same event, and they don't carry the same commercial weight.
A mention just means the brand's name showed up somewhere in the answer. It might be accurate. It might be outdated, buried at the bottom of a long list, or framed negatively, and the mention counter won't know the difference.
A citation means the system displayed a source link. That doesn't mean the linked page actually shaped what the AI said. A cited page might have supplied one minor fact while the model's pretrained knowledge, or an uncited source entirely, drove the actual recommendation.
A recommendation is the AI explicitly telling the user "go with this brand." It's the clearest commercial signal of the three, and it's also the rarest.
There's a fourth layer that's mostly invisible to tracking tools: answer influence, where the brand's content shaped the response even when the brand's name never appears. Surface Labs made this point in a June 2026 framing: a report that presents single-run outputs as settled truth can genuinely mislead an executive who trusts the number at face value.
Blend all four signals into one visibility percentage and you get a number that can rise even as recommendation share, the thing that actually matters commercially, falls. A useful GEO report keeps mention rate, citation rate, recommendation rate, and sentiment as four separate lines, because each one calls for a different response from the team reading it.
The platform variance problem: a brand's score on one AI engine says almost nothing about its score on another
Each major AI platform runs its own retrieval index, its own citation logic, its own source weighting. Each major AI platform runs its own retrieval index, its own citation logic, and its own source weighting, and that lack of shared architecture across them is why there's no reason to expect a brand's performance to transfer from one to the next.
ChatGPT pulls from a blended retrieval stack that includes Bing, OpenAI's own web index, and several third-party providers. It rewards broad web authority and editorial press coverage, and it tends to paraphrase without naming sources at all, unless web search has been explicitly triggered.
Perplexity works differently: live web retrieval, explicit numbered citations on every response, and more source slots per response than ChatGPT offers. That means less competition per citation slot. Reddit alone accounts for close to half of Perplexity's top citation sources.
Gemini draws heavily from Google's own index and from YouTube, so brands with strong Google authority and a real video presence tend to do better there than elsewhere.
Claude has been noted for relatively traceable citations in 2026 comparisons, which makes it a stronger fit for YMYL categories (your-money-or-your-life topics like health and finance). Analysis of a large citation dataset found Claude gave brands the highest owned-domain citation share of any platform measured.
Cross-platform overlap is thin. Cross-platform citation overlap is thin, research found only around 11% of domains cited by ChatGPT also appear in Perplexity results for identical queries, and a substantial share of queries produce completely disjoint source sets across platforms. Cross-platform analysis has documented extreme citation volume variance for the same brand, depending purely on which platform was asked.
Average all of that into one composite "AI visibility score" and the number hides exactly the information a brand needs: which platform it's winning on, which one it's losing on, and why those two things require completely different fixes.
Most tooling doesn't even attempt to cover a blind spot that this produces. Monitoring tools are built around horizontal assistants, ChatGPT, Gemini, Claude, Perplexity, Copilot, because they can be prompted from outside. Retailer-owned vertical agents, Amazon's Alexa for Shopping (formerly Rufus) and Walmart's Sparky among them, sit much closer to the actual transaction, and almost no tool tracks both categories at once. Visibility in one says close to nothing about visibility in the other.
Why synthetic query testing and real user behavior diverge in ways that matter for brand strategy
Most AI visibility platforms run synthetic queries through developer APIs and organize whatever comes back into a dashboard. The queries themselves are chosen by the vendor. They aren't derived from what real users actually type.
That matters because the API environment isn't the consumer app. The consumer product carries memory, personalization, location context, its own model-routing logic, and its own rules about when to bother searching the web at all. None of that is visible in a bare API call.
Real prompt monitoring looks different. Consented consumer panels that track real app and browser behavior reveal commercial intent signals that synthetic testing cannot reproduce, because synthetic testing never observes an actual person with an actual goal.
The scale gap is stark too. Real user query sets run into the millions. Most tracking platforms test dozens to a few hundred prompt permutations, which is exactly why the IAB's August 2026 framework set 50 queries as a floor for even a directional reading.
Profound's research across more than 27 million real answer engine citations found that 97.4% of AI citations come from non-Tier-1 earned media: Reddit threads, niche YouTube videos, LinkedIn posts, small vertical sites. Synthetic query sets built around obvious, branded prompts tend to miss exactly the sources that drive citations in the open-ended, messy way people actually ask questions.
The sound approach runs both tracks side by side. Synthetic testing shows how an engine responds to a controlled stimulus. Panels show what people actually ask and what sources turn up in those real answers. Neither one alone tells the full story.
What score volatility looks like in practice and why it makes trend lines unreliable
Research shows that only about 30% of brands visible in one AI response appear again in the very next response to the identical query. A brand can be "visible" in one run and gone in the next without a single thing changing on its end, no new content, no competitor move, nothing.
A month-over-month trend line built from daily single-query snapshots is a line drawn through noise. Each individual point is already shaky on its own, and the slope connecting two shaky points carries even less real information than the points themselves.
Model updates create false inflection points on top of that. A platform-wide shift in citation behavior can produce a sharp, sudden drop in reported scores, and the dashboard will present that drop as a brand performance event. The actual cause never appears inside the number.
That 8%-to-11% swing mentioned earlier isn't a special case. A move of that size between reporting periods falls squarely inside the variance range that repeated-run research predicts, and it can't be pinned on a brand action, a content change, or a competitor's move without a lot more evidence than a dashboard usually provides.
The arXiv paper's core finding bears repeating here: visibility should be described as a probability distribution. A brand might turn up in a large majority of responses to a given prompt category, or only a small fraction, and that range, not any single point inside it, is the meaningful output. A workable minimum protocol runs 30 to 50 queries, each fired 3 to 5 times each, across the platforms being tracked. That produces a directional reading. Anything short of that is one observation wearing a metric's clothes.
How to read AI visibility scores without being misled by them
Jason Goldberg's framing in Forbes, August 2026, cuts right to it: these numbers are unreliable enough that they'd struggle to survive a basic statistics review, and yet brands waiting for cleaner measurement before they start tracking will end up a year or two behind the ones who started anyway. The right move is building the practice now while pushing the methodology to improve alongside it, not waiting for a perfect number that isn't coming.
Report visibility as a range, never a point. "The brand appeared in roughly 60% to 75% of responses across this prompt set" carries real information. "11.3%" carries the illusion of precision and not much else.
Split the four signal types before anything reaches a stakeholder. Mention rate, citation rate, recommendation rate, and sentiment each point to a different problem and a different fix, and folding them into one number erases exactly the detail a marketing team needs to act on.
Track each platform on its own line. A composite score averaging ChatGPT and Perplexity hides the one piece of strategic information that determines where the brand is winning, where it's losing, and what's different about those two platforms that explains the gap.
Anchor every trend claim to how many runs produced it. A score change only means something if each data point came from enough repeated queries, 30 to 50 at minimum, run more than once. Anything built from single snapshots should get labeled indicative, not definitive, before it reaches an executive audience.
Mark model update dates directly on the chart. A sudden jump or drop that lines up with a known model release or index change isn't a brand event, and it shouldn't get presented as one.
Pair the synthetic tracking with some behavioral signal that lives outside the dashboard: AI referral traffic (thin as it is, given that 93% of AI search sessions end without a click), brand search lift, direct traffic trends. These act as an independent check on whether the visibility score actually correlates with real discovery behavior, or whether it's just measuring itself.
And track across platforms simultaneously, not one at a time. ChatGPT, Perplexity, Gemini, Claude, Copilot, and whatever comes after them each run different retrieval logic and produce different citation patterns, and a tool built for multi-LLM monitoring from the ground up gives coverage that a single-engine reading, or a hybrid SEO tool stretched to cover one more channel, simply can't replicate. That's the difference between a real measure of brand presence and a number that only looks like one.
Sources
- AI Visibility Numbers Are Unreliable. Measure Them Anyway
- Don't Measure Once: Measuring Visibility in AI Search (GEO)
- AI Visibility Measurement Is Messy: How to Report GEO Without Fake Precision | Surface Labs
- 9 AI Visibility Optimization Platforms Ranked by AEO Score (2026)
- Why GEO Scores Aren’t the Solution To Your Brand’s AI Engine Visibility
- AI Visibility Scores Shift Between Runs
- searchenginejournal.com
- authoritytech.io