Measuring AI Visibility More Often Makes It Worse
Past a certain point, extra measurement stops adding information and starts manufacturing movement. The fix is more repetitions per point, not more points.
Measuring more often makes the data worse. Not more expensive, not merely redundant: actively worse, in the specific sense that it produces changes to explain that were never there. It is the most common way teams get this wrong, and it is the opposite of the instinct everybody arrives with.
The reason sits in the mechanism. Language models are probabilistic. Ask the same question twice, in two fresh sessions, and you can get two different lists of recommended suppliers, with nothing at all having changed in the world, on your site, or in the model. That variance is not an error to be eliminated. It is a property of the thing you are measuring, and every decision about cadence follows from taking it seriously.
A single run is not a measurement
If one run can name you and the next can omit you, then a check is a coin, and a check today compared against a check last week is two coins compared against each other. Whatever difference you find, you can construct a story for it. Marketing shipped a page. A competitor published something. The model changed. All of these are plausible, none of them is supported, and the exercise is astrology with a dashboard.
What carries information is a proportion: across many runs of the same prompt, how often were you named. That number is stable enough to compare. A single answer is an anecdote about one sample from a distribution.
Which reframes the whole cadence question. The instinct is to ask how frequently to measure. The useful question is how many repetitions a single measurement needs before it means anything. Get that wrong and increasing the frequency just gives you more unreliable points to draw a line through, and a line through noise looks exactly like a trend.
How many repetitions
There is no universal number, and anyone quoting one has not looked at their own data. But the shape is knowable and you can find yours in an afternoon.
Take one prompt that matters. Run it 10 times, in 10 fresh sessions, and write down whether your brand appeared. Then look at the sequence. If you were named in 9 of the 10, or in none of them, the prompt is stable and a small number of repetitions will do. If you land somewhere in the middle, that prompt sits on a boundary, and a single run of it is worthless: it will flip, and every flip will look like news.
Boundary prompts are the ones that generate false alarms, and they are also usually the interesting ones, because they are where you are close to being included. Identify them and treat them differently: more repetitions, and a rule that you do not report movement on them at all unless the proportion moves substantially across a whole measurement cycle.
The budget is fixed, so this is a trade
Repetitions are not free, and this is where the abstract argument turns into arithmetic you have to actually do.
Every measurement is a prompt, on an assistant, in a language, repeated some number of times. Multiply those out and the total grows quickly. A modest set of 20 prompts, across 4 assistants, in 5 languages, is already 400 answers for a single repetition of everything, and one repetition of everything is precisely the thing we established tells you nothing.
So you are spending a fixed budget across four dimensions that all want more of it: more prompts, more assistants, more languages, more repetitions. Something has to give, and it should be a deliberate choice rather than whatever the default was.
The plans reflect that trade rather than hiding it. On PSentry the entry plan monitors 20 prompts on a single site and scans 30 prompt runs a month; the next steps up move to 100 and 300 and 600 monitored prompts, with 150, 500 and 1200 scannable runs a month respectively. Scans run twice a month on every plan. Those numbers are worth reading as a constraint rather than as a feature list: they tell you that you cannot simultaneously have a wide prompt set, every language, and heavy repetition, and that pretending otherwise just means the repetition silently goes to one.
My preference, for what it is worth, is to cut the prompt set before cutting the repetitions. A narrow set of questions you can actually trust beats a broad set of numbers that each flip at random. Twenty prompts you believe is a better instrument than two hundred you do not.
Match the cadence to how fast the thing actually changes
The second half of the argument is about the underlying process rather than the sampling.
What determines whether an assistant names you is, roughly, what the web says about you and what a training run absorbed. Neither of those moves on a daily scale. Third-party coverage accumulates over months. Documentation gets written and then sits there. A training run happens when a provider decides it happens. There is no daily process underneath, so daily sampling is sampling a monthly signal sixty times and calling the difference movement.
Twice a month is a reasonable cadence against that, and that is why the product runs on it rather than offering something faster. It is also why real-time alerting does not exist here and is not on the way. An alert that fires when a single run stops naming you would fire constantly, on variance, and after the third false alarm nobody reads it. Selling that as a feature is easy and it would be a way of charging for noise.
The one moment when a step change is real
There is an exception, and knowing it is most of what separates reading these numbers well from reading them badly.
Providers ship new model versions. When they do, answers can shift sharply and immediately, and that shift is genuine rather than sampling noise. The way to tell the two apart is to look at whether the movement is concentrated or general.
If a handful of prompts moved and the rest of the set sat still, something specific happened: a page of yours, a competitor's content, a change in how one question gets interpreted. If nearly everything moved at once, on one assistant, in the same direction, you are almost certainly looking at a new model version rather than at anything you or your competitors did. Those two findings lead to completely different responses, and a team that only tracks one aggregate score cannot distinguish them, because the aggregate looks identical either way.
This is a good argument for keeping the per-prompt detail available rather than collapsing everything into a single headline number. The headline is for the board. The diagnosis lives underneath it.
A cadence that survives contact with a real quarter
Putting it together, without pretending it is more precise than it is.
Fix the prompt set and stop editing it. A set you keep rewording produces movement that is entirely your own doing, and it destroys the comparison you were measuring for. Add prompts when you enter a market or launch a line. Do not reword the existing ones because a better phrasing occurred to you on a Tuesday.
Measure on a schedule rather than when someone asks. Twice a month is enough, and the schedule matters more than the frequency because it is what makes the points comparable.
Report by quarter, not by cycle. Two measurement points cannot establish a direction, whatever the arrow on the dashboard suggests.
And when a number moves, check first whether the whole set moved. That single question resolves more false alarms than any other, and it costs nothing.
Record enough on every run that the comparison is still possible in six months, because the version of you that has to explain a change will not remember. For each answer: the date, the assistant, the language, the prompt exactly as it was sent, whether your brand was named, which competitors were named, and whether the answer cited any sources or was written from memory. That last field is the one people leave out and later wish they had, because an answer that retrieved and an answer that did not are two different measurements wearing the same number. Everything else is reconstructable. Those are not.
The uncomfortable summary is that most of the value here comes from measuring a little, carefully, for a long time. It is a worse story than a live dashboard. It has the advantage of being about something.
Before you set a cadence
We check manually every morning. What is wrong with that?
The check itself is fine as a habit. The problem is treating consecutive mornings as a series. One run per morning is one sample per day from a distribution that varies within the day, so the day-to-day differences you see are mostly the distribution, not your position changing.
Is there any reason to measure more than twice a month?
Around a specific event, yes, and only if you spend the extra budget on repetitions rather than on more prompts. A product launch or a site migration is a case where you want a dense before-and-after on a small set of questions. Outside that, extra frequency mostly buys variance.
How many runs before I believe a change is real?
Enough that the proportion, not the individual answer, has moved, and enough that it has held across more than one measurement cycle. If a prompt named you in most runs and now names you in most runs minus one, nothing has happened.
Different assistants disagree with each other. Which one is right?
None of them, in the sense you mean. They read different sources and retrieve differently, so disagreement between assistants is a finding about the assistants rather than an error to resolve. Track them separately and do not average them into one figure.