Back to blog How to Read an AI Visibility Scan, and What to Do With It

How to Read an AI Visibility Scan, and What to Do With It

The score is the least useful thing in your first report. Here is what a scan actually does, what it can tell you, and what no tool can promise.

A scan finishes, a report appears, and the first thing almost everyone does is look for the score. They decide it is good or bad, and close the tab. That is the least useful thing you can do with the report.

What follows is how a scan actually works: what happens between the moment you add a site and the moment a number appears on a dashboard, what the output contains, and what it is honestly good for. This is the one article on this blog that talks openly about the product, so it is also the one that has to be most careful. A measurement you do not understand is a measurement you cannot challenge, and you should be able to challenge this one.

The cycle, step by step

You define the site and the markets

You give the system a domain and tell it which languages and markets you actually sell in. This sounds trivial, and it is the step people get wrong most often: they list the markets where they have a website rather than the markets where they have revenue. Those are different lists, and the gap between them is one of the more interesting things a scan surfaces later.

A prompt set is generated for the site

The prompts are generated automatically, per site, from what your site says it does. The point is not to ask who is [your brand]. A model will always find something to say about a named brand, so a branded prompt mostly measures whether you exist, not whether you are recommended.

The prompts that matter are the ones a buyer asks before they know you exist. Who supplies this category in Europe. What is the best option for a company of this size. Compare the leading vendors for this problem. Those answers have room for a handful of names, and either you are one of them or you are not. That is the thing being measured.

The prompts run across four platforms

ChatGPT, Claude, Gemini and Perplexity, in each of your languages. Four, not one, because they do not behave alike. Perplexity retrieves and cites live pages. The others lean more heavily on what they absorbed during training, which means they can describe you fluently and never link to you, or describe a version of you that is badly out of date. A brand can be present in one and invisible in another, and averaging that away would hide the only part of it that is actionable.

The answers get read, not just counted

Each response is parsed for what it says about you and about everyone else in the answer: whether you appear at all, how you are described, which sources are cited, and which competitors are named in the same breath or in your place. That competitor list comes from the answers themselves, not from a list you supply, which matters: models cluster brands by how the web talks about them, not by how your market map is drawn.

A score is computed per language and per market

Not one global score. One per language and market, because visibility does not travel. A brand can be recommended confidently in English and absent from the identical question asked in German, for the boring reason that the German-language web says less about it. A single global figure would average those two facts into a number that describes neither.

Why the scans are scheduled, and why that is the right call

Scans run on a schedule: twice a month. We would rather explain that than let you assume it is a limitation we have not got around to fixing.

A brand's standing inside these models moves slowly. It changes when the web changes: when documentation gets published, when a comparison article gets written, when a forum thread accumulates answers, when a model is retrained or its retrieval index refreshes. None of that happens on a Tuesday afternoon. Meanwhile the models are probabilistic: ask the same question twice in the same hour and you can get two different lists of vendors, with nothing underneath having changed at all.

Put those two facts together and continuous monitoring produces mostly noise. You would watch a line wobble, attribute the wobbles to your own actions, and act on a signal that was never there. A scheduled cadence, repeated across many prompts, is what turns individual answers into a pattern. Real time would be a marketing claim: better on a feature list, and less informative.

What the report actually contains

  • Presence. Whether you appear at all in answers to the prompts your buyers would plausibly ask.
  • Mentions. Where you appear, and in what terms. Being named as an option and being named as an afterthought are both mentions, and they are not the same thing.
  • Citations and sources. Which pages the models leaned on to say what they said. Your own domain sometimes. More often, somebody else's.
  • Competitors named in your place. The brands that occupy the slot you wanted, market by market.
  • Comparison across languages and markets. The same prompt set, the same platforms, different answers.
  • The export gap. The markets where you sell, or want to, set against the markets where an AI actually names you. These two maps rarely overlap as neatly as anyone expects.
  • Trend over time. How each market moves across successive scans, which is the only view in which a single number starts to mean something.

How to read the first report

Not as a report card. As a map of where to look.

The score is the least informative thing on the page the first time you see it, for a simple reason: you have nothing to compare it against. A score is a baseline before it is a judgement, and it only starts working on its second reading.

The genuinely useful output of a first scan is the list of sources the models use to talk about your category. That list describes where authority currently sits in your market, as the machines understand it. It will contain some names you expected and, usually, several you did not: a trade publication you ignore, a comparison site you have never submitted to, a forum where your product is discussed by people who have never spoken to you, a competitor's documentation quoted to explain a problem you also solve. It is not a to-do list handed to you by a tool. It is evidence, and what you do with it is a commercial judgement.

The second most useful output is the competitor list: it tells you who the models think you are comparable to. The third is the gap between your home market and your export markets, normally wider than anyone in the room is prepared for, and the one finding that reliably changes a plan.

What it does not do, stated plainly

None of this is PSentry reaching into the models on your behalf. What actually happens is narrower: four systems get asked, their answers get read and laid out plainly, and that is where the tool's job ends, well short of anything that could be called engineering an outcome. No result comes with it, because producing one was never something a reporting tool could hand over in the first place.

What you get is narrower and more honest: a repeated, per-market picture of what four AI systems say when someone asks about your category and does not mention your name. Whether that is worth anything depends on whether you are prepared to act on it. The tool tells you where you stand. It cannot make you move.

There is a public API and CSV export, for an unglamorous reason: the data is yours, and you can pull it into your own reporting or leave and take your history with you. Look at PSentry and judge it on its own terms, which is the only sensible way to evaluate any kind of ai visibility monitoring.

Frequently Asked Questions

My score went down between two scans. Did I do something wrong?

Possibly not. Model outputs vary between runs, so a small movement in either direction can be probabilistic drift rather than a real change in standing. This is exactly why a trend across several scans means something and a single delta between two of them usually does not. Look for direction sustained over time, not for a story that explains the wobble.

Why not scan more often, if a scan is just an API call?

Because frequency would not improve the signal. The thing being measured moves on the timescale of the web changing and models being retrained, not on the timescale of a dashboard refresh. Sampling a slow, noisy variable more often gives you more noise, not more information.

What if the report says I am invisible everywhere?

Then you have learned something most companies in your position do not know at all, and you have learned it before it costs you a market. Absence is a finding, not a failure of the measurement. The source list from that same scan is where you would start.

Does being cited mean I will get traffic?

No, and the two should stay separate in your head. A model can describe your product accurately and send you nobody, and it can send you a buyer who never clicked anything. Which is why an analytics tool cannot answer this question, and why the measurement has to happen in the answers themselves.

Is this a replacement for our SEO reporting?

No. It is a second scoreboard, measuring inclusion in generated answers rather than placement in a list of links. The two share plumbing, since both depend on what is crawlable, but they are not the same measurement and one does not substitute for the other.