In short
- 86 to 94% gone
- That many ChatGPT citations vanished between February and April 2026 across five markets — because of two changes at OpenAI, not because of anything a website did. (seoClarity)
- And then it came back
- In May the citations recovered to nearly their pre-March level. The researchers thereupon revised their own conclusion: not a decline, but volatility.
- 5 · 56 · 74
- The share of sources that changes weekly in AI Overviews, Google AI Mode and ChatGPT Search. So a single figure for “stability” does not exist. (SISTRIX, 82,619 prompts)
- Core or carousel
- For 86% of prompts there is a stable core of one to five sources; everything around it rotates at 89% per week. The question is not whether you are in the answer, but in which part.
What does the study say?
Between February and May 2026 seoClarity followed millions of ChatGPT questions across five markets. On 8 March and 19 April OpenAI changed something, and the number of citations fell by 86 to 94%. In Germany 85% of questions at one point got no source at all. Then, in May, it largely came back.
SISTRIX looked at a different question: how often do the sources in an AI answer change from week to week? Across 82,619 prompts, seventeen weeks and six countries. The outcome depends entirely on the platform — AI Overviews swaps 5% per week, Google AI Mode 56%, ChatGPT Search 74% — and on the level you look at: 74% at domain level, 85% at page level. The study itself warns against adding those figures together.
| What was measured | How | What the study does not say |
|---|---|---|
| seoClarity — 5 markets, February to May 2026, millions of prompts | tracked citations per prompt, and counted the share of prompts with no source separately | nothing about the quality of an individual site — it is a platform change, and the researchers say so themselves |
| SISTRIX — 82,619 prompts, 1,548,213 snapshots, 17 weeks, 3 platforms, 6 countries | compared weekly which domains are new and which remain, at domain level | that a single stability figure exists — the study explicitly warns against lumping platforms together |
seoClarity (Mitul Gandhi, May 2026, updated June 2026) · SISTRIX (Johannes Beus, 1 May 2026). Both figures were found back in the source text itself on 10 August 2026.
Where can you check this yourself?
-
seoClarity — Mitul Gandhi May 19, 2026, updated June 19, 2026
five markets (US, UK, Canada, Germany, Italy), February to May 2026, millions of measured prompts
Between February and April 2026 the number of ChatGPT citations fell by 86 to 94% across all five markets, after two platform changes on 8 March and 19 April. In May it largely came back. The researchers revised their own conclusion: not a decline, but volatility.
The collapse is a PLATFORM change and says nothing about the quality of an individual site — the researchers state this explicitly themselves. Anyone who made a change in March would wrongly conclude it failed; anyone who acted in May, wrongly that it worked.
-
SISTRIX — Johannes Beus May 1, 2026
82,619 prompts, 1,548,213 snapshots, 17 weeks (17 Dec 2025 – 8 Apr 2026), 3 platforms, 6 countries
How volatile AI sources are depends entirely on which platform you look at: AI Overviews swaps 5% of its sources per week, Google AI Mode 56%, ChatGPT Search 74%. For 86% of prompts there is a stable core of one to five domains; everything around it rotates at 89% per week.
The study says itself that the figures are CONSERVATIVE: at URL level the churn is 85% instead of 74%. And it explicitly warns against lumping platforms together — “lumping the platforms together obscures more than it reveals”.
Every link goes to the original publication. What you read above was found back in that text itself on August 10, 2026 — not taken over from someone else's summary.
Why does this matter to you?
Because it decides whether someone can sell you a result they did not produce. Suppose in March 2026 you paid an agency to improve your AI visibility. In April someone looked: your citations had collapsed. Conclusion: it did not work. In reality OpenAI had changed something twice and everyone had dropped. Turn it around and it gets worse: whoever looked in May saw everything recover and could defend any invoice without having changed a single letter on your site.
And the second study adds a second trap. If someone tells you AI sources are “stable” or indeed “volatile”, the first question is: on which platform, and at which level? The same study yields 5% and 74% depending on which engine you look at. An average across three platforms is not a summary but a blend.
What does ceeme do with it?
After a fix we measure the same questions on the same engines again, and put before next to after. Without that re-measurement we make no statement about results — that is not caution but the only position these two studies allow. If something drops after an intervention, we roll it back.
And we measure each engine separately rather than turning it into one figure. For the same reason: an average across platforms that churn at 5% and 74% hides exactly what you would base a decision on.
Frequently asked questions
Does this mean a fix is pointless?
No, the opposite. It means you cannot judge a fix on a single measurement afterwards. The intervention can work perfectly well while the figure drops, and vice versa. What you need is the same measurement before and after, across several engines, so that a platform change shows up as something hitting everyone at once.
Why cite research that qualifies your own story?
Because the nuance is the argument. If this page only said AI answers are volatile, we would be leaving out half of the same measurement — in AI Overviews, 53% of prompts see not a single source change across seventeen weeks. That difference is exactly why you have to measure per engine instead of believing one figure.
How do I know whether my site is in the core or the carousel?
By measuring the same question repeatedly over a longer period and seeing whether your domain is there every time or on and off. That is exactly what a monitoring series does, and it is something other than a snapshot — a single measurement cannot by definition tell core from carousel.
Read on
Free, no account and no card.