Skip to content
ceeme

Can you measure a position in an AI answer?

No. Three independent studies from 2026 reach the same finding: asked the same question, an AI assistant almost never returns the same brand list twice, and hardly ever in the same order. A position number therefore measures not your visibility but the coin toss of that one moment. What you can measure is how often you appear across many repeats, with the margin of error attached — and that margin decides whether a shift means anything.

Measure your own company → Rather talk it through first

In short

Fewer than 1 in 100
The odds that the same question gives the same brand list twice. In the same order: fewer than 1 in 1,000. (SparkToro, 2,961 runs)
30 points of noise
That is how far a visibility figure can swing at 21 repeats without anything having changed. (PetraLabs, 7,200 runs in one day)
19% wrong
That is how often a single measurement points the other way than twelve measurements of the same question. So a fifth of your report is wrong. (MaxAEO, 23,040 answers)
So what does work
Breadth beats depth. At the same budget, 200 prompts × 3 repeats reaches ±5.4 points, against ±15.4 points for 20 prompts × 30 repeats.

What does the study say?

Three teams, three set-ups, the same outcome. SparkToro had 600 volunteers put the same twelve questions to ChatGPT, Claude and Google AI nearly three thousand times: the odds of the same list twice stayed under 1 in 100. PetraLabs ran six buyer questions 600 times per engine in a single day, to rule out day-to-day effects, and measured how far a figure swings from chance alone. MaxAEO gathered 23,040 answers across six engines and looked at how often one measurement disagrees with twelve.

What they jointly show is not a defect but a property. A language model does not rank brands; it draws from a pool of names it finds fitting for that question, and reshuffles that pool with every answer. The size of that pool and how often you are in it are real quantities. The spot you held that one time is not.

What was measuredHowWhat the study does not say
SparkToro — 2,961 runs, 600 volunteers, 12 prompts, 3 enginesthe same question repeated, logged out, counting how often the list recursthat visibility is unmeasurable — only that a POSITION is
PetraLabs — 7,200 runs, 6 prompts, 2 engines, one daybootstrap over 10,000 draws to set the threshold below which a difference is chancenothing about what happens over weeks — all measured in one day, deliberately
MaxAEO — 23,040 answers, 18 brands, 320 prompts, 6 engines12 repeats per prompt across fourteen days, one run compared with the majoritymeasured on B2B software questions; a local question with two known names needs less depth — the study says so itself

SparkToro (Fishkin & O'Donnell, 2026) · PetraLabs (Tom Zu, June 2026) · MaxAEO (Chris Han, July 2026). Every figure on this page was found back in the source text itself on 10 August 2026.

Where can you check this yourself?

  • SparkToro — Rand Fishkin & Patrick O'Donnell (Gumshoe.ai) 2026

    600 volunteers, 12 identical prompts, 2,961 runs across ChatGPT, Claude and Google AI

    The odds of getting the same brand list twice are under 1 in 100. In the same order: under 1 in 1,000.

    The study does not say visibility is unmeasurable — it says a POSITION is. Presence across many runs remains statistically measurable.

  • PetraLabs — Tom Zu June 2026

    7,200 runs in a single day — 600 repeats per prompt per engine, six buyer questions, ChatGPT and Gemini, logged out

    At 21 runs, visibility can swing up to 30 percentage points from chance alone. At 50 runs, a real 10-point shift is detected only 17% of the time.

    All measured on one day, to rule out day-to-day effects. What happens over weeks — a model update — therefore falls outside this measurement.

  • MaxAEO — Chris Han July 21, 2026

    23,040 answers — 18 brands, 320 prompts, 6 engines, 12 repeats per prompt across fourteen days

    A single run disagrees with the twelve-run verdict 19% of the time. And at the same budget, 200 prompts × 3 runs reaches a margin of error of ±5.4 points, against ±15.4 points for 20 prompts × 30 runs.

    Measured on B2B software questions, where the lists are long. For a local question with two or three known names, the results sit closer to 0% or 100% and need less depth — the study says so itself.

Every link goes to the original publication. What you read above was found back in that text itself on August 10, 2026 — not taken over from someone else's summary.

Why does this matter to you?

Because it decides whether you can defend an invoice. If a report moves you from fourth place to second, the honest question is not *why* but *how many times was it measured*. At twenty-one repeats that same figure can travel thirty points without anything having happened. Selling such a shift as a result is selling noise with a story wrapped around it.

What does ceeme do with it?

We show no position. We show whether you are mentioned, cited or chosen, across several repeats, with the answer the engine actually gave beside it — so you can see for yourself where the figure comes from. And we choose breadth over depth: rather more real customer questions than the same question thirty times, because at the same budget that is demonstrably more accurate.

That costs us something, and rightly so. A position number sells more easily than a range, and a dashboard with arrows looks more decisive than one saying *this difference is too small to mean anything*. We show the second, because three independent measurements say the first does not exist.

Frequently asked questions

Does this mean AI visibility cannot be measured?

No, and the distinction matters. What is unmeasurable is the POSITION. Presence measures fine: how often you appear across many repeats, in what wording, and with which sources attached. That is a percentage with a margin of error, not a ranking.

Other tools do show a position. Are they wrong?

They show something that exists — in THAT answer you were in that spot. The question is whether it repeats, and on that three measurements point the same way. We do not judge someone else's product; we say what we show and why, and put the sources underneath so you can check.

So how many repeats do you do?

That depends on your plan, and it lives on the pricing page rather than here — a number you buy does not belong in a study. What does belong here is the choice underneath it: at a fixed budget we pick more questions over more repeats, and the figure that choice rests on is above.

Is there research that contradicts this?

On the volatility of cited SOURCES, yes — measurements there differ sharply, and that is a different question from this one. As soon as we have checked those measurements at the source, a separate page will lay them side by side, including the one that proves us wrong. What we will not do is show the half that suits us.

Read on

Free, no account and no card.