How to choose an AI visibility tool: the criteria that matter
No verdicts and no comparison table: the questions that separate a tool that measures from one selling you noise, and the sales answers that should worry you.
A lot of products will now show you a dashboard with your brand's AI visibility on it: a score, a trend line, an arrow. They look alike from the outside, and that is the problem. A number on a screen looks the same whether it came from a careful measurement or from one lucky query.
So this is not a review: no winner at the end, no comparison table. Just what separates a tool that measures from one that produces plausible output, and how to check it yourself. Some of these questions are awkward to put to a salesperson. Ask them anyway.
What the thing is supposed to do
An ai visibility tool has one job: tell you whether generative AI systems name your brand when someone asks the questions your buyers ask, and what they say when they do. That is harder than it sounds, because language models are probabilistic. The same question can return two answers with two different sets of brands. There is no stable position to read off the way you read a search ranking, only a distribution, and a distribution can be estimated only by sampling it repeatedly.
How many systems, and which ones
ChatGPT, Claude, Gemini and Perplexity are not four skins on the same engine. They were trained differently, they retrieve differently, and they hold different opinions about your market. A brand recommended confidently on one can be absent from another, and you cannot predict which without looking.
How to check. Ask which platforms are covered, then ask the follow-up almost nobody asks: which model, specifically. A vendor who can name it knows what their system is doing.
How many times each prompt is run
This never appears in a feature comparison, because it is boring and it is where the cost sits. It matters more than anything else here. A tool that runs each prompt once is reporting a coin flip: you appeared, or you did not, and tomorrow it may flip with nothing about your brand having changed. Stack a month of single runs into a trend line and the chart moves convincingly while meaning nothing, because what you are watching is the model's own variance rendered as a business metric. Sampling the same prompt several times turns an anecdote into a frequency: not "the model mentioned us" but "the model mentioned us in most of the answers it gave". Only the second is worth acting on.
How to check. Ask how many executions per prompt, per scan. If the answer is one, you know what the trend line is made of. If it arrives as a description of the methodology rather than as a number, treat that as one.
Which languages, and how
Visibility does not travel across languages. A model can only repeat what exists, and the German or Spanish web usually says far less about a company than the English one does, so a brand recommended in English can be absent from the same question asked elsewhere. If you export, a tool that measures only in English is measuring the market where you least need help: the blind spot is the market you have never visited, where a model is quietly recommending a local competitor in your place.
How to check. Ask whether prompts run natively in each language or get translated and run in English behind the scenes: those are different measurements. Then ask whether results are split per language, because an average across markets hides what you are looking for.
Mention and citation are not the same event
A model can describe your product accurately, at length, and never link to you. Being named without a link is a brand outcome and sends no traffic; being cited is a traffic outcome. A tool that collapses both into "we found you" has discarded the distinction that decides what you do next.
How to check. Open a raw answer where the brand is discussed and see whether the tool records that it carried a link back to your site. If the interface cannot express the difference, the data underneath probably cannot.
Does it show you the sources
When a model retrieves, it pulls from somewhere, and knowing where is the closest thing to an actionable output this category produces. If the systems talking about your market keep quoting one directory and a competitor's comparison page, that is where the conversation about you is happening, and it is not on your own site.
How to check. Ask to see one full model answer with its sources, exactly as it came back. A tool that gives you a score but never the evidence behind it is handing you a verdict you cannot audit.
Transparency of method
Nearly every tool here, ours included, computes some kind of visibility score. A score is useful shorthand, and a liability the moment you cannot see how it was built: you cannot tell whether it moved because your visibility changed or because the vendor changed the weighting, and you cannot compare it with anyone else's.
How to check. Ask to see the exact prompts being run for your brand, and whether you can change them. Prompts are the instrument: if you cannot see them, you do not know what is being measured. Then ask what goes into the score, in what proportion. "Proprietary" translates as: trust the number, do not check it.
Can you take the data with you
The raw answers, the mentions, the sources, the dates: that is your record of how the machines described your brand over time. A tool that will only ever show it inside its own dashboard is selling you a subscription to your own history. Check for a CSV export and an API before you sign, not after.
The answers that should worry you
Some claims are not exaggerations, they are tells.
- "We will increase your AI visibility." A monitoring tool observes. It has no mechanism to change what a model says about you, because there is no lever to pull: no submission form, no index, no ranking factor to tune.
- "We optimize AI responses." Nobody outside the labs that build these models optimizes their responses. What you can influence is what the web says about you, which is slow, indirect, and not what that sentence is selling.
- "We guarantee you will appear in the answers." Nobody can guarantee an outcome from a probabilistic system they do not control. That is not a stretch of the truth, it is simply false.
- "Real time alerts." Models do not change their mind about a brand between breakfast and lunch. A tool that pings you constantly is selling urgency rather than information.
- "GEO" that turns out to mean maps and local listings. Generative Engine Optimization is about generative engines. If the pitch drifts into proximity and business profiles, the vendor has confused two disciplines that share an acronym.
A serious vendor will say the deflating sentence out loud: we measure, we do not move the needle for you. Anyone unwilling to say it is either confused about what they built, or hoping that you are.
Now apply all of it to us
Full disclosure, and it should be obvious by now: we build one of these. PSentry runs prompt sets across ChatGPT, Claude, Gemini and Perplexity, in each language a brand sells in, and reports where the brand is named, where it is cited, which sources the models draw on, and which competitors get recommended in its place. We have a conflict of interest in publishing this page, and you should read it with that in mind.
Which is the point. Put the questions above to us. Ask how many times we execute each prompt, which model sits behind each platform name, what our score is made of, and tell us if you find it too opaque: that criticism applies to us as much as to anyone. Two things we will say before you ask. We already said the deflating sentence once: we measure, we do not move the needle. Twice a month is simply how often we check, not a design flaw waiting for a real-time upgrade nobody actually needs. And touching what a model outputs is not a button we have, not a partnership we can buy into, not anything money changes: ours is a report, not a remote control.
Then pick whichever tool survives the questions. A category this young does not need better marketing. It needs buyers who ask how the number was made.
Frequently Asked Questions
How many prompts should a tool be running for my brand?
Fewer than you would guess, if they are the right ones, and each run several times. A long list executed once per scan is worse than a short list sampled properly: breadth without repetition just collects noise from more directions.
Can I check this by hand instead of buying anything?
Yes, and you should, at least once. Ask each platform what a prospect would ask, in every language you sell in, and read who gets named. What you cannot do by hand is repeat it consistently over time, which is what turns an impression into a trend.
Are visibility scores comparable between tools?
Generally not. Different prompts, platforms, sampling and weightings mean two scores are not the same quantity wearing different names. Compare the inputs instead: when two tools disagree, the useful question is which prompts each ran.
Is any of this actionable, or am I buying a thermometer?
It is closer to a thermometer than most vendors admit, and that is not an insult. The actionable part is not the score: it is the sources the models use, and the competitors they name instead of you. Those point at places on the web where the story about your market is being written without you in it.