What an AI Visibility Score Measures, and What It Cannot
A score compresses hundreds of AI answers into one number. Here is what survives the compression, and what quietly gets lost along the way.
Somewhere on a dashboard there is a number with your brand's name beside it, and this month it went up. The next question is the only one that matters: up because of what? A score will not tell you. It was never built to.
That is not an attack on scores. It is a description of what they are. An AI visibility score takes a pile of messy evidence, many answers, to many prompts, across several assistants and several languages, and squeezes it into one figure you can read in a second. The squeezing is the point. But squeezing is lossy, and it is worth knowing what gets left behind.
A score is an index, not the work
Think of it the way you think of a stock index. The index tells you the market fell. It does not tell you which company had a bad quarter, or why. Nobody manages a portfolio by staring at the index alone.
What a score is genuinely good at is direction. It compresses hundreds of model answers into something that can be plotted, compared, and put in front of someone who will never read a raw model output. It answers "is this getting better or worse, and where do I look first". Real questions, both. Simply not the same question as "why".
The number is the index of the work, not the work. The work sits underneath it, in the answers themselves, and a tool that hides those answers from you is asking you to trust an aggregation you cannot audit.
What actually goes into it
Tools build their scores differently, but the honest components are variations on the same handful of observations. Each means something specific, and they get confused with one another constantly.
Presence: does the answer name you at all
The most primitive signal, and the most brutal. You asked a model which suppliers it would recommend for something you sell. Your name appeared, or it did not. No page two, no near miss, no "almost ranked". Presence is binary inside a single answer, which is why it is useless measured once and meaningful measured many times.
Mention frequency: how often, across repeated runs of the same prompt
The component most people skip, and the one that turns presence into information. The same prompt, run again, can return a different set of names. So the interesting quantity is not "were you named" but "in how many runs of this prompt were you named". A brand that appears in nearly every run holds a stable place in the model's answer space. A brand that appears occasionally sits at the edge of it. Ask once, and the two look identical.
Citation: being named is not being linked
A model can describe your product accurately, at length, and never send anyone to your site. A citation, an actual link to one of your pages offered as a source, is a different event from a mention, and it behaves differently by platform. Perplexity retrieves and shows its sources inline. A model answering from what it absorbed in training has nothing to point at, because it is not reading anything at that moment. Merging the two into one bucket makes a score easier to compute and harder to interpret.
Share: who is standing next to you in the sentence
An answer that recommends vendors almost always recommends more than one. Your share of the names in that sentence is a more honest measure of position than your presence in it. Being named beside a couple of rivals is not the same as being named in a long list, and being the brand that gets named while a rival is left out is the only form of ranking that survives here. This component also tells you who your competitors are, which is often not who you think.
Coverage: how much of your prompt set you appear in
Presence on prompts you already win is comfortable and uninformative. Coverage asks the harder question: of all the buying questions you decided were worth monitoring, in how many do you appear at all? The gaps are the only thing on a dashboard that suggests what to do next.
Why one number is a dangerous thing to own
Language models are probabilistic. That sounds like a technicality. It is not. It means the same prompt, asked twice, can return two different lists of vendors, and neither answer is broken. Sampling is part of how these systems produce text.
The consequence is uncomfortable for anyone who wants a clean metric. A score computed from a handful of runs is mostly measuring noise. It will move between scans for reasons that have nothing to do with you: a different sampling path, a different retrieval set, a model update nobody announced. Treat that movement as signal and you will spend a quarter explaining a fluctuation that was never real. Which is why the absolute value of a score is close to meaningless, and two things about it are not:
- The trend. Measured the same way, on the same prompts, over time. A number that drifts in one direction across repeated scans is telling you something. A number that jumps once is telling you about variance.
- The comparison between segments, above all between languages. Visibility does not travel. A brand the models discuss confidently in English can be effectively invisible in the same question asked in German, because the German-language web says less about it. Your home market against an export market is a like-for-like comparison, and the distance between them is usually the most actionable thing on the screen.
This is why PSentry computes its AI Visibility Score™ per language and per market rather than as one global figure: a single worldwide number averages away the differences you most need to see.
What a score cannot tell you
It does not predict traffic
You can be cited and receive no click. The entire design of an assistant answer is to spare the user the trip. Someone can read a paragraph that names your product, form an opinion, and never touch your domain. Visibility in AI answers and sessions in your analytics are different quantities, and expecting one to move the other on a predictable schedule is how people end up disappointed by a metric that was working fine.
It does not measure the quality of the mention
The largest blind spot in any single number. Being named as the obvious choice for serious buyers and being named as the cheap alternative for people watching their budget both count as one mention. So does being named in a sentence that gets your product category slightly wrong. A score records that you were in the room. It does not record how you were introduced, and the introduction is frequently the part you would most want to change. There is no substitute for reading the answers.
It is not comparable across tools
Every tool in this category picks its own prompt set, its own platforms, its own repetition per prompt, its own weighting of the components above. Two scores built from different ingredients are not two measurements of the same thing. They are two different instruments. Pick a method, keep it stable, read its trend.
The honest version
A high score is not a business objective. Nobody was ever paid for a score. It is a thermometer: it gives you the temperature, it tells you when the temperature is moving, and it tells you which room to walk into. Everything that actually matters, why a market went cold, why a rival is named instead of you, whether the models even describe your product correctly, lives in the answers underneath the number and not in the number itself.
Use the score to decide where to look. Then go and look.
Frequently Asked Questions
What is a good AI visibility score?
There is no universal threshold, and anyone quoting you one is describing their own scale rather than a standard. The only meaningful benchmarks are yourself over time, and yourself in one market against yourself in another.
My score dropped between two scans. Should I be worried?
Not yet. A single move can easily be sampling variance rather than a real change in how models treat your brand. Check whether the direction persists across several scans, and whether it shows up in more than one language. A drop visible everywhere at once is more likely to be real.
Why do two tools give me different scores for the same brand?
Because they measure different things and call them by the same name: different prompts, different platforms, different repetition, different weighting. Neither the score nor the trend is portable between them.
Is being cited better than being mentioned?
They do different jobs. A citation can send a reader to your site. A mention with no link still shapes what a buyer believes before they ever run a search. Track both, and do not let a tool quietly merge them.
Can I raise my score directly?
Not by touching the score, and not by touching the model behind it. What moves the number is what the web says about you, in each language, and a monitoring tool only reads that back to you, every couple of weeks rather than in real time. It does not manipulate model outputs, and nobody watching your score on your behalf can promise it will climb.
How often is it worth recomputing?
Less often than instinct suggests. Models do not revise their view of a brand overnight, and the web underneath them moves slowly too. A steady cadence produces a readable trend. Refreshing constantly produces noise you then feel obliged to explain.