Skip to content
ceeme

Why do two tools give me a different AI score?

An AI visibility score summarises how often you appear in the answers of a set of engines, weighted by how heavily each engine counts in your market. Two scores for the same company therefore almost always differ, and the cause rarely lies in the measurement itself: it lies in which questions were asked, which engines were counted, and with which market share those engines were weighted.

By Njusja Orban, founder of ceeme4 min read

Measure your own company → Rather talk it through first

In short

Three dials
Which questions, which engines, which market. Turn one of them and the same company gets a different figure.
Weighting
An engine barely used in your country should not count as heavily as the largest one. So we weight per market.
Coverage
What percentage of your market the measured engines cover together. Without it you do not know what the score is worth.

Why does not every engine weigh the same?

Because the balance differs sharply per country and an unweighted average lets the smallest engine count as heavily as the largest. Below are the actual shares for our four markets, with the month — a figure without its month is a figure without an expiry date.

Note what these figures measure: referrals to websites, so which assistant sends people to a site. That is different from how many people use an app, and for this question it is the right number — we measure visibility that leads to a visit.

MarketMonthLargestThe rest, in order
BelgiumJuly 2026ChatGPT 73.80%Gemini 10.42% · Copilot 6.96% · Perplexity 4.51% · Claude 4.22%
NetherlandsJuly 2026ChatGPT 77.48%Gemini 8.67% · Copilot 4.83% · Claude 4.60% · Perplexity 4.33%
FranceJuly 2026ChatGPT 75.96%Gemini 10.03% · Perplexity 4.96% · Claude 4.92% · Copilot 3.92%
GermanyJuly 2026ChatGPT 73.07%Gemini 11.38% · Perplexity 6.44% · Claude 4.73% · Copilot 4.34%

Statcounter · AI chatbot market share — referrals to websites, not user counts. Read per country on gs.statcounter.com; the month is given per row because countries are not published at the same pace.

What happens to an engine we cannot measure?

It counts in the denominator and not in the numerator, or it visibly falls outside the weighting. What does not happen is that we add the market up to a hundred per cent over only the engines we happened to measure — then coverage looks complete while a third is missing.

An engine that did not actually search during a measurement also gets its own label and does not count as zero. Zero means "we looked and you were not there"; not searched means "we do not know", and mixing those two turns a technical failure into a poor score.

How do I tell whether two scores are comparable?

By three things that must accompany both: the question list, the engine list, and the market with the month of the weighting. If one is missing you are comparing two figures without knowing whether they are about the same thing — and then the gap between them means nothing.

When is a score simply worthless?

When your own name was in the questions. Then you measure whether the engine can read what is written about you, and that is almost always yes. Such a score heads towards ninety and says nothing about the customer who does not know your name yet.

And when it was measured once. An assistant's answer is not identical every time; one measurement on one day is a snapshot. What has meaning is asking the same set again and looking at the difference — including the times nothing changed.

Frequently asked questions

Which score is good?

That question has no general answer, and any provider giving one is selling a feeling. What does have meaning is the comparison with yourself three months ago, same questions and same engines. A threshold that holds across sectors does not exist: in a category with three providers forty per cent is a lot, in one with two hundred it is exceptional.

Can I check the weighting myself?

Yes, and that is the point of the table above. The shares come from a public source you can look up per country, and the month is stated. A weighting you cannot recalculate is a weighting you have to trust — and trust is exactly what a measurement should make unnecessary.

Read on

Free, no account and no card.