AI Visibility Scoring Methodologies Across Leading Platforms
Different AI platforms score brand visibility in incompatible ways.

AI Visibility Score, or AVS, is the metric marketers use to measure whether a brand shows up inside AI-generated answers, and no two platforms calculate it the same way. Each platform surfaces brands differently, and treating them as interchangeable is where most measurement frameworks break down. Anyone treating "AI visibility" as one score, tracked one way, across every platform, is measuring the wrong thing.
What an AI Visibility Score measures and how it is calculated
An AI Visibility Score is a normalized number, usually 0 to 100, that rolls up six separate signals into one directional read on where a brand stands. Those six signals are mention rate (how often the brand shows up across a defined set of prompts), presence quality (is the brand the primary recommendation, a secondary mention, or just a name dropped in passing), platform breadth (strong in ChatGPT but invisible in Gemini is a real gap, not a rounding error), sentiment framing (the tone the model uses when it talks about you), share of voice relative to competitors, and citation rate, meaning how often the platform actually links back to a source.
The mechanics, as laid out by Campaign Creators in May 2026, start with 20 to 30 unbranded prompts built around actual buyer questions, not searches for the brand name itself. Each prompt runs across multiple AI engines, and each response gets scored 0 to 5: a 5 means the brand was the primary recommendation, 3 is a secondary mention, 1 is a passing reference, and 0 means the brand never showed up. Running 20 prompts across 4 platforms produces 80 scoring events, with a maximum possible raw score of 400. A raw score of 136 works out to an AVS of 34.
A companion metric worth knowing is Share of Model, or SoM, credited to Jack Smyth and built out further by Tom Roach. It's the AI-era version of share of voice, calculated as brand mentions divided by total category mentions across the tracked prompt set, times 100. SoM is earned through brand mentions relative to total category mentions, unlike old-school share of voice, which is bought with media spend. And it behaves differently because the game itself is smaller: a typical LLM response cites somewhere between two and seven domains, compared to the ten organic slots Google used to hand out on a results page. That makes AI citation a tighter, more zero-sum contest than traditional search ever was. What counts as a "good" score varies by category too. A mature, competitive vertical will have a different visibility ceiling than a niche one, so benchmarks only mean something when read against others in the same industry.
Why visibility scores are directional signals, not precise verdicts
Only a small fraction of brands, 16% according to McKinsey, systematically track AI search performance at all right now. The discipline is young, and the tooling shows it.
Scores also move around more than most dashboards let on. Research cited by alhena.ai found that roughly 30% of brands visible in one AI response for a given query don't reappear in the very next response for that same query. That's the baseline condition of the medium, not noise to be ignored. Four things drive that fluctuation: how a prompt is worded, which model version answers it, how fresh the underlying index is, and the plain stochastic randomness baked into how these models sample and generate text.
The clearest case study here is the ChatGPT 5.0 update in September 2025. After the rollout, outbound citations across the platform dropped, and brand tracking dashboards everywhere showed sharp declines that had nothing to do with content quality or marketing execution. The model had simply changed how, and how often, it named its sources. Anyone reading that dip as a verdict on their brand's relevance was reading it wrong.
A single week's Share of Model reading is a snapshot, not a scorecard. Trend lines across weeks, aggregated across platforms, tell the real story. Point-in-time numbers, no matter how confidently a vendor dashboard presents them, should be treated as directional. That's a known limitation even among the more careful measurement providers in this space, so it's better to build a reporting cadence around it rather than fight it.
How ChatGPT scores and surfaces brands and drives most AI referral traffic
ChatGPT is the biggest single surface in this category by a wide margin. Similarweb put it at 68% of the AI chatbot market as of January 2026, OpenAI reported 900 million weekly users in the first quarter of that year, and the platform processes over a billion prompts a day. Scale like that means ChatGPT drives the overwhelming majority of AI referral traffic, 87.4% of it across the ten industries studied in Conductor's 2026 benchmarks.
Standard ChatGPT tends to paraphrase rather than cite. A brand can be the answer, worded right into the response, without a single clickable link pointing back to where that information came from. The default experience most users get does not include clickable sources in the way dedicated search modes can provide. This makes ChatGPT the largest citation surface in existence and, at the same time, the hardest one to trace. Referral link analysis won't catch most of what's happening. Structured, repeated prompt testing is the only way to see it.
Recency matters a lot here. Freshness carries real weight on ChatGPT, particularly for news, tech reviews, and any category that moves fast. Then there's the October 2025 entity update, which changed how ChatGPT recognizes and recommends brands, likely groundwork for commerce features inside the platform down the line. Brands trying to optimize for that shift are building one clear primary entity per page, backed by three to six supporting entities tied to Wikipedia, WikiData, and pillar content that reinforces the connection.
OpenAI also launched advertising inside ChatGPT on February 6, 2026, in the form of clearly labeled "Sponsored" cards that appear below the answer itself. Ads don't touch the model's actual response. Getting recommended is still something a brand earns through the underlying data, not something it buys. For scoring purposes, that means ChatGPT's mention rate is the single most consequential number in this whole exercise by volume, but it has to be captured through prompt testing rather than referral analytics, because referral analytics simply can't see most of what's going on inside the platform.
Perplexity's Citation Explicitness and Source Transparency
Perplexity does the opposite of ChatGPT by default: every response comes with clickable, numbered source links laid out for the user to see. Citation isn't inferred here, it's printed on the page. That makes Perplexity the most traceable of the major platforms, by a clear margin.
Profound's AEO guide reports that Perplexity handles over 100 million queries a week, a scale smaller than ChatGPT's but still substantial. A study from GOYBO International Journal of Marketing Intelligence, authored by Iyappan, rated Perplexity's citation explicitness as "Very High," the top rating among the five major platforms profiled. Recency weighting also came in "Very High" in that same study, which lines up with Perplexity's real-time web index. Freshness isn't a tiebreaker here, it's close to a primary ranking factor. Structured data sensitivity rated high too: content built with clear headings, comparison tables, and FAQ sections gets rewarded in ways that look a lot like classic SEO best practice, just applied to a different surface.
Because the citations are explicit and clickable, Perplexity-sourced traffic appears more cleanly in analytics tools than ChatGPT-sourced traffic does. That makes a brand's Perplexity citation rate a useful diagnostic well beyond Perplexity itself. If a content or structure change moves the needle there, it's a decent early signal that the same change might help on other platforms too, even before the data comes in from those slower-moving surfaces. Comparison tables, step-by-step guides, expert Q&A formats, and answer-first content blocks are the formats that consistently win here, because they're built in a shape the model can lift and present directly.
Google Gemini and AI Overviews: Where SEO Authority and AI Citation Overlap Most Directly
Gemini does double duty. It runs as a standalone conversational assistant, and it also powers Google's AI Overviews, the generated summaries that now sit directly inside regular search results pages. That second role is what makes Gemini the bridge between the SEO world marketers already know and the newer discipline of AI citation.
The scale is enormous. Conductor's 2026 benchmarks looked at 21.9 million Google searches and found 5.5 million of them triggered an AI Overview, which is 25.11% of all searches in the sample. But that number swings hard by industry. Healthcare queries triggered an AIO 48.75% of the time, the highest of any sector studied. Financials came in at 25.79%, Utilities at 25.4%. Real Estate sat at just 4.48%, the lowest in the dataset, with Consumer Staples close behind at 6.82%.
The cost of not showing up in that summary is steep. Seer Interactive studied informational queries and found organic click-through for the top-ranking page fell 61% once an AI Overview appeared, from 1.76% down to 0.61%. That's a page that used to rank first, still ranking first, and getting a fraction of the clicks it used to get.
Gemini's grounding in Google's existing search infrastructure means that established SEO signals, structured data, entity recognition, and content authority, remain relevant here in ways that carry over from traditional search. Plenty of GEO advice treats LLMS.txt files as a real lever for Gemini visibility, though that claim remains contested and unconfirmed by Google. Plenty of GEO advice out there still treats it as a real lever, even though it isn't one.
One finding from Conductor's dataset makes the authority point sharply. In the Financials industry, NerdWallet captures 6.73% of AI citations, more than the traditional banks it covers. Being a bank doesn't buy citation share. Being the publisher banks and consumers both treat as a trusted reference does. And structured data has a measurable payoff on Gemini specifically: a 2025 Relixir study of 2,100 pages found FAQPage schema correlated with a substantially higher probability of AI citation.
Claude's Authority Signals, Source Weighting, and Structured Data Sensitivity
Claude is less transparent about citations than Perplexity but still surfaces sources more than standard ChatGPT in many contexts. The 2026 Iyappan study rated its citation explicitness "Moderate," meaning it's less consistently transparent about sources than Perplexity, which labels every citation explicitly. Recency weighting rated "Low to Moderate," lower than either ChatGPT or Perplexity, which actually works in favor of evergreen, well-sourced content that isn't tied to a news cycle. Structured data sensitivity landed "Moderate to High": organization helps, but Claude leans more on who's saying something than on how neatly it's formatted.
That "who's saying it" question is the whole ballgame on Claude. Across AI search broadly, roughly 85% of brand mentions trace back to third-party pages rather than the brand's own site, and AirOps' analysis of more than a billion citations found that brands are several times more likely to get cited through a third party than through their own domain. That pattern is most visible on Claude. Industry publications, analyst reports, consumer review sites, trade press, and earned media coverage are what the model actually pulls from, not homepages or product pages a brand controls directly.
An SE Ranking analysis of 129,000 domains found that third-party brand web mentions are the single strongest predictor of AI citation, carrying a 35% weight in the model they tested, and that pattern is especially relevant on a platform like Claude where third-party sourcing is central. Practically, that turns PR and earned media into a genuine GEO lever. External validation from outlets a model already treats as credible is how Claude tells a brand with real standing apart from one that just publishes a lot of content about itself. Long-form, deeply sourced material with named experts and visible author credentials tends to outperform the short, answer-first blocks that do so well on Perplexity. Different platform, different format, different win condition.
Conductor's Cross-Platform Benchmark Data on Industry-Level Citation Patterns
Conductor's 2026 AEO/GEO Benchmarks Report is the closest thing this space has to an industry-wide baseline. It covers 13,770 domains across ten major industries, drawing on billions of sessions and more than 100 million AI citations, which makes it the first dataset of its kind sized to say something real about how citation behavior differs by sector.
AI referral traffic still makes up a small slice of total web visits overall, 1.08% across the ten industries measured, though it's climbing at roughly 1% month over month. IT led every industry studied at 2.8% of total traffic coming from AI referrals, with Consumer Staples next at 1.9%. And across all ten industries, ChatGPT alone accounted for 87.4% of that referral traffic, reinforcing just how much of this ecosystem runs through one platform even though that platform is the hardest to measure directly.
Domain-level citation share tells its own story, industry by industry. In Consumer Staples, Amazon holds 17.99% of domain citations. In IT, Google captures 5.34%. In Financials, it's NerdWallet at 6.73%, ahead of the banks themselves. Real Estate is the odd one out: Zillow holds 7.36% of brand mention share despite not cracking the top five most-cited domains in the category, which shows that being mentioned by name and being cited as a source aren't always the same fight.
Category authority decides who wins AI citation, not company size and not ad budget, and that pattern holds consistently across the data. Category authority decides who wins AI citation, not company size and not ad budget. A brand with a strong third-party footment in its category, one that gets written about, reviewed, and referenced by other trusted sources, will outperform a bigger competitor that spends more but has thinner earned coverage. That's the thing to build toward, more than any single platform's scoring quirks.

