How to Test Any AI Visibility Tool Yourself, Including Ours
No comparison table, no verdicts on products we have never used. A protocol you can run in a week, plus the checks that expose a weak tool. Apply it to us.
You probably landed here looking for a comparison table. A grid with the competitors down one side, the features across the top, green ticks in our column and grey dashes in theirs. We are not going to give you one, and it is worth explaining why before you decide whether to keep reading.
We do not use our competitors' products. We have not run their prompt sets, we have not seen their dashboards from the inside, we do not know what their plans cost this month or what their scans actually do. Anything we wrote about them would be reconstructed from their marketing pages and our own imagination, and shaped, inevitably, by the fact that we would rather you bought from us. A vendor reviewing a product it has never held is writing fiction with a footnote, and readers can tell. Articles exactly like that used to sit on this blog. We took them down.
So here is the honest version. Instead of telling you who wins, we will hand you the method we would use ourselves to test two tools against each other, and then state what our product measures and what it refuses to claim. Run the method on us. If we fail it, do not buy us.
The week-long protocol
You can settle this yourself in about a week. Most tools in this category offer a trial. The trick is not to look at them one after the other, forming vague impressions. It is to run a single experiment through both of them at once.
Build one prompt set, not two
Write the questions your buyer would actually type before they know you exist. Not "what is [your brand]", which tells you nothing, because the model will always find something to say about a name you hand it. The useful prompts are the ones where you are not mentioned: "best tool for X for a mid-sized company", "who supplies Y in Europe", "alternatives to [the market leader in your category]". Cover your main use cases, then freeze the list. The same list goes into both tools, word for word.
Freeze everything else too
Same languages in both. Same markets. Same competitors named, if the tool lets you name them. Same window of time, running in parallel, not one this month and one the next, because the web moves and so do the models. If both tools are in trial simultaneously and fed identical inputs, any difference in the output is a difference in the tool. That is the entire point, and it is destroyed the moment you let one variable float.
Read the outputs side by side
At the end of the window, do not compare the headline scores. Every vendor computes its score differently, so putting them next to each other is meaningless theatre. Compare the underlying claims instead: which prompts produced a mention, on which platform, in which language, and who else was named. That is the layer where a tool is either right or wrong.
The checks that separate a measurement from a rumour
Now the part that actually decides it. Five questions, to be asked of every vendor in the running, including us.
Take three claimed mentions and verify them by hand
Pick three cases where the tool says you were mentioned. Open the platform yourself, ask the same question, and look. You will not get a character-for-character match, because these models are probabilistic and no two runs are identical. That is not what you are checking. You are checking whether you are plausibly in that answer at all, whether the competitors named resemble the ones the tool reported. If a claimed mention is nowhere to be found across several attempts, ask the vendor to show you the stored response text. If they cannot produce it, the mention was never verifiable, and neither is anything else in that dashboard.
Ask how many times each prompt is run
This is the question that quietly disqualifies the most tools. Language models do not return the same answer twice. Ask which vendors lead your category, ask again an hour later, and you can get a different cast of characters. A tool that runs each prompt once and reports the outcome as a fact is selling you a coin flip with a chart around it. Ask directly: how many executions sit behind a single reported result, and is the score an aggregate or a snapshot? Vague answers are answers.
Ask whether it distinguishes a mention from a citation
These are different events with different consequences. A model can describe your product accurately, at length, and never link to you. Another can link to your documentation while recommending someone else. One is brand presence, the other is a referral path. A tool that folds both into a single number has thrown away the distinction that would have told you what to do next.
Ask to see the sources
When a platform retrieves live pages before answering, those pages are the reason you were included or ignored. A tool that captures them lets you discover that a comparison article you have never heard of is quietly deciding your visibility in Spanish. A tool that does not gives you a score and no explanation, which means no action either.
Ask whether you can take your data out
Export is not a convenience feature, it is a test of confidence. A vendor happy to hand you the raw rows expects them to survive inspection. A vendor who will only show you data inside their own charts has decided that the presentation matters more than the substance. Ask for an export during the trial, not after you have signed.
What PSentry measures
Now hold us to the same standard. Here is the product, without adjectives.
PSentry runs your prompt set across four platforms: ChatGPT, Claude, Gemini and Perplexity. It does this in each language and market you sell in, and reports where you were named, where you were cited, which sources the models used, and which competitors were recommended in your place. It produces a visibility score per language, so you can see the gap between the market where you are known and the ones where you are exporting into silence. You can export the data.
That is the whole product. A measurement instrument.
What PSentry does not do
This list matters more than the one above, because it is the list nobody else volunteers.
- It is not a live feed. The scans behind it run a couple of times each month, and that pace is deliberate: we are not going to bolt on instant alerts just to manufacture urgency around a number that moves this slowly.
- It does not do geolocation. The segmentation is by language and market, not by where a user is physically standing. If you want local search visibility, that is a different discipline and we are not it.
- It cannot be gamed from our side. There is no dial here that makes a model talk about you more favorably, and if another vendor in this trial claims otherwise, that claim is exactly the kind this article told you to go verify.
- No outcome is on offer. Whether a model warms to you afterward depends on what the rest of the web says about you, not on anything a subscription can arrange, so treat what you get here as a reading, never as a result we owe you.
- It does not plug into your other systems. No chat integrations, no CRM sync, no business intelligence connectors. There is an API and there is CSV.
We are a young company. We have no decade of accumulated wisdom to sell you, and this category is new enough that anyone claiming otherwise is embellishing.
Then test us
Put us in the same trial window as whatever else you are considering. Identical prompt set, same languages, same weeks. Take three mentions we claim and check them in ChatGPT yourself. Ask how many times we execute a prompt. Ask to see the sources. Ask for your data in a file, and see how quickly you get it.
If we do not hold up, buy the other one. We would rather lose a customer to a tool that measured honestly than keep one who discovers, months in, that the number they have been reporting upward was noise.
Frequently Asked Questions
Why not just publish a comparison and disclose the bias?
Because disclosure does not repair the underlying problem: the facts would still be wrong. Competitor pricing, features and limits change constantly, and a vendor reconstructing them from a rival's landing page gets them wrong in ways that happen to be convenient. A disclaimer does not make invented details true.
The tools give me different scores for the same brand. Which one is right?
Possibly neither, and the question is malformed. Scores are proprietary compositions of different inputs, so they are not comparable across vendors. What is comparable is the evidence underneath: the prompts, the responses, the mentions, the sources. Judge the evidence and let the scores be whatever they are.
What if a tool refuses to answer these questions?
That is a result, and you should record it as one. Every question on this page has a straightforward answer that a vendor knows about its own system. Reluctance to give it is not a scheduling problem.
Do I need a tool at all?
Not necessarily. If you sell in one language and care about a handful of questions, open the platforms and ask them yourself, on a repeating reminder. Tools become necessary when the matrix of prompts, languages, markets and platforms grows past what a person can run by hand without quietly giving up.