In short
- Variation
- The same question, a different answer. Not because anything changed about you, but because the model is not deterministic.
- Snapshot
- What one measurement gives you. Useful as a starting point, dangerous as proof — especially if someone builds a success claim on it.
- Rotating measurement
- Repeating and rotating the question set, so you take not just chance but also one-sidedness out of your number.
- Re-measurement
- The same questions, the same method, after a fix. Without that, an improvement is an assertion.
What does the study say?
The core is not that AI is unreliable, but that a measurement of AI has a spread — as does any measurement of something non-deterministic. Ignore that and you are measuring noise and calling it a result.
| What was measured | How | What the study does not say |
|---|---|---|
| brand visibility in AI answers, measured repeatedly | analysis of the variation between runs of the same measurement | how many measurements are enough for your sector — that depends on how strongly your questions fluctuate |
(GEO research) (2026). Don't Measure Once: Measuring Visibility in AI Search. arXiv 2604.07585
Why does this matter to you?
Because it decides whether your number is worth anything. A tool that measures once and then claims a rise cannot tell that rise apart from the spread of the measurement itself. That is not a small nuance: it is the difference between a result and a fluke.
What does ceeme do with it?
Our monitoring re-measures on rotation and shows the trend over time rather than one snapshot. And after a fix we measure the same prompts again with the same method — if something drops, we roll back. This research is the reason we claim no result without a re-measurement.
Frequently asked questions
Why not just measure once?
Because you would not know whether you are seeing something or chance. One measurement is a fine starting point for roughly where you stand; it is no basis for reading a change from.
How often do you measure, then?
That depends on the route. Monitoring re-measures continuously and on rotation; after a fix we re-measure the same prompts. The frequency belongs to the route and not to this study, which only says that once is too few.
Does this mean AI visibility cannot be steered?
No, the opposite. Precisely because there is noise, it is worth measuring with a method that takes the noise out — otherwise you are steering on fluctuations. What you cannot do is guarantee a position, and that is a different matter.
Read on
Free, no account and no card.