Prompt monitoring: which questions reveal your AI visibility
In AI search the unit of measurement is not the keyword, it is the prompt. Pick the wrong ones and you measure nothing, pleasantly.
In search, the unit of measurement was the keyword. You picked a phrase, checked where you ranked, and the number moved or it did not. Every tool and every report was built on that atom.
Generative engines took the atom away. Nobody types a keyword at an assistant. People type a situation, in sentences, with context, and get back one answer naming a few companies. So the thing you measure is now a prompt. Your score, your trend, your competitor list: all of it is a function of which prompts you chose. Picking the wrong ones is the easiest way to measure nothing and feel reassured while doing it.
A test that everyone passes is not a test
Throw out the prompts that contain your own name. "What is [brand]?" "Is [brand] any good?" With your name in the prompt you have handed the model the entity, and the hard part, deciding who is relevant enough to bring up, is already done for it. It will always find something to say and it will sound informed, for a market leader and for a company nobody has heard of alike. A test that cannot fail carries no information.
Those prompts confirm. They do not measure. Their one legitimate job, which belongs in its own column, is checking whether the description is accurate. That is a factual audit, not visibility.
Five shapes of a prompt that measures
Every prompt worth logging shares one property: your name does not appear in it. That is what turns a question into a measurement, because it reproduces the situation that matters commercially: the buyer does not yet know you exist. Within it, prompts fall into five families, and a set missing one has a hole in it.
Categorical
"What is the best inventory system for a mid-sized food producer?" Someone is building a shortlist out of nothing. The model returns a handful of names. You are in them or you are not: no partial credit.
Comparative
"Compare the main options for warehouse automation." This forces an explicit set, which exposes what a categorical prompt hides: the brand the model treats as the default, the one everything else is measured against. Sometimes it is not on your competitor slide at all.
Problem-shaped
"Our returns keep piling up and customers are complaining, what do we do?" No category, no vendors, just the pain. The model picks the category first and the vendors second, so there are two ways to lose: absent from the list, or filed under a category you are not in, where the question of naming you never arises.
Substitution
"What can I use instead of a traditional [category]?" This catches a buyer who is already leaving something. Being named here is a different signal from appearing in a category list: you have been filed as a viable replacement, which is closer to a sale.
Qualifying
"Which option suits a regulated industry with almost no internal IT?" This is where deals get scoped, and it is the family everyone forgets. A model that names you for large enterprises and never for small teams has said something precise no aggregate score will surface.
Write like a buyer, not like a keyword
"Best CRM" is a keyword. Nobody says it to an assistant. What they say is closer to: "we are a mid-sized manufacturer selling across Europe, our CRM does not talk to our ERP and sales refuses to use it, what should we look at?"
Context is not decoration. Models condition heavily on it: company size, sector, budget, timeline, an integration that has to work. Each of these changes the names that come back, and a prompt stripped of them measures a buyer who does not exist. Use your customers' vocabulary, not your positioning deck's: a prompt written in language nobody outside your company says will answer confidently about a market that is not yours.
Once is an anecdote
These models are probabilistic. Ask the identical question twice and you can get two different lists, in a different order, with a different brand on top. That is not a defect to engineer around, it is the medium. So every prompt has to be executed repeatedly, on each of the four major assistants, which disagree with each other far more than any disagrees with itself.
What you are measuring is not presence but frequency: named in most runs, named in some, never named. Three different commercial positions, and one run cannot tell them apart. The real size of a set is prompts multiplied by languages, by platforms, by repetitions, so when you have to cut, cut prompts, never repetitions.
What to record on every run
Most people log a yes or a no, the least informative field there is. For each execution, capture:
- Named. Did your brand appear at all.
- Cited. Was there an actual link, or only a mention.
- Named in your place. Every brand in the answer, whether or not you are among them. The most useful thing generative engines hand you, and it is free.
- Order. First in the list or last. Not a ranking, but it tracks how confidently the answer treats you.
- Characterization. Being cited as "the affordable option" is not the same data point as "the standard choice for large operations", though both count as a mention. Everybody skips this column and it carries more information than the rest combined.
- Platform, language, date. Without these, nothing above can be compared to anything.
Do not move the ruler
Once a set works, freeze it. A visibility number only means something relative to the prompts it was computed over: the moment you edit the set you have changed the instrument, and any trend line crossing that edit is fiction.
Sets do change eventually: the category picks up new vocabulary, you enter a market you were not in. Handle it like a schema migration, not a document edit. Version the set, log what was added or removed and why, and keep the old prompts running alongside the new for a while, so you can see what the change itself did. Rarely, and never quietly.
The trap: a set built to make you look good
Here is how the exercise gets ruined, almost never deliberately. Someone writes a prompt, runs it, does not appear, dislikes that, rephrases, appears, keeps the second version. Do it a dozen times and you have a set with a flattering score attached that measures one thing: your own ego.
The defence is procedural. Write the whole set before you run any of it. Take them from places where you cannot cheat: sales calls, support tickets, the first email a prospect sends. Then run it as written and accept the result. If the set contains nothing you expect to lose, it is not an instrument. Appearing everywhere is not evidence that you are winning. It is evidence that the set is wrong.
Where this stops being doable by hand
A modest set, in every language you sell in, across four platforms, repeated often enough to mean anything, is more executions than a person can read and classify. That is where prompt monitoring has to be automated, and it is what PSentry does: it generates a prompt set per site and runs it across ChatGPT, Claude, Gemini and Perplexity, recording who gets named in your place.
With a caveat we would rather say out loud. A generated set is a starting point, not an oracle: it has never sat in one of your sales calls. If a prompt does not sound like something a real buyer of yours would type, it is not measuring your market. And the boundary stays where it is: this is measurement. It reports what the models said. It does not change what they say.
A set that never says your name, sounds like a person talking, runs many times instead of once and holds still long enough to be compared against itself will tell you where you stand. One built from questions you already know you win will tell you nothing, pleasantly, for as long as you pay for it.
Frequently Asked Questions
Should I put competitor names in my prompts?
Yes, in the comparative and substitution families, because real buyers do exactly that. A competitor's name is not the same mistake as your own: it anchors the question without handing you an appearance. Keep those prompts a minority of the set, or you end up measuring their position instead of yours.
Can I just translate my prompt set into other languages?
Translation is not enough. Buyers elsewhere phrase the problem differently, the category may not carry the same name, and the incumbent they compare everything against is often a local company you have never heard of. Localize the set, then treat each language as its own instrument.
Can I ask an AI to write the prompt set for me?
It is a reasonable way to produce raw material. What comes back is how a model imagines your buyers speak, which is smoother and more generic than how they actually do. Treat it as a first draft and check it against language you have heard from a real customer.
What does it mean if I am never named, in any prompt?
Two very different diagnoses hide behind the same result: either the models have no usable impression of you, or they have one and do not consider you a fit for the framings you are testing. The "named in your place" and "characterization" fields are what separate them, and the response differs in each case.