In short
- Pinned, never "the latest"
- A name meaning "latest version" swaps under your measurement without anything reporting it.
- Every measurement carries its stamp
- What it ran on is recorded, so afterwards you can see whether two measurements are comparable.
- A fallback is not free
- It keeps the measurement alive and it starts a new series. That is stated rather than buried.
Which model versions is measuring done on today?
Below, per engine, with the fallback alongside. That second column exists because a provider does not only replace a version number but sometimes makes it unreachable: the call then returns an error and you measure nothing, which in a chart looks like a score of zero.
One engine in that table has no fallback, and it says so. An empty cell would read as "not applicable"; it is a risk you should know about before starting a series on that engine.
| AI search engine | What we measure on today | If that version drops away |
|---|---|---|
| ChatGPT | gpt-5.6-terra | gpt-5.4-mini |
| Gemini | gemini-3.5-flash | gemini-3.6-flash |
| Perplexity | sonar-pro | sonar |
| Claude | claude-sonnet-5 | claude-sonnet-4-6 |
| Grok | grok-4.5 | no fallback — the series stops and starts over |
| Mistral | mistral-medium-latest | mistral-small-latest |
| DeepSeek | deepseek-v4-flash | deepseek-chat |
| Meta AI | Llama-4-Maverick-17B-128E-Instruct-FP8 | no fallback — the series stops and starts over |
What happens to my series if a model is retired?
It stops and a new one begins, and that is reported rather than measured through. The temptation to simply carry on is strong, because the chart then looks continuous — and that is precisely the more dangerous of the two: an unbroken line across two different models looks like a trend and is not one.
Why does nobody notice that something like this happened?
Because no error appears. A replaced model simply answers, only slightly differently: other sources, another order, another number of names in the list. That is the same kind of error as valid structured data saying the wrong thing — everything works, and the outcome is wrong.
Can I keep an eye on this myself without a tool?
Partly, and it is worth doing. Write down per measurement which question you asked, on which day and in which assistant, and keep the answer itself rather than your summary of it. That is exactly what a series does, by hand, and it is enough to see whether a shift came from you or from them.
What you will not do by hand is keep it up across four engines, in two languages, every week. That is where the difference sits, not in the idea.
Frequently asked questions
Do you measure on the free or the paid model of an assistant?
On the version in the table above, and it is there with its number so you can check. What a visitor sees in the free version of the same assistant may differ — that is a real limitation and it is here because nobody publishes it.
Do you add a new engine as soon as it exists?
Only if it is used in your market. Adding an engine that is barely queried in your country raises the number of measurements and not their meaning — and it dilutes your average with an audience you do not have.
Read on
Free, no account and no card.