Skip to content
ceeme

Who checks the AI that writes your copy?

A second model, from a different family, that is not allowed to know who wrote the text. Three studies from 2025 and 2026 show why both requirements are needed: a model judging its own work marks its own failures as passed more than 50% more often — even on criteria that can be verified by machine — and the name attached to a text shifts the verdict by up to fifty percentage points, even when that name is wrong. Blind evaluation removes the second; only a different model family does anything about the first.

Measure your own company → Rather talk it through first

In short

More than 50% more often
That is how much more often a judge wrongly marks a criterion as satisfied when the text is its own. Measured on rubrics a machine can check — so there is nothing to argue about on taste. (Pombal, Rei & Martins)
A smarter model does not help
Across twenty models there is no link between capability and less bias, and sometimes it runs the other way. What does help is a structured rubric: 31.5% less on average. (Yang et al.)
The name shifts the verdict
And not always upwards: one label raised the scores, another lowered them — up to fifty percentage points of difference in preference, even when the name was wrong. (Saraf et al.)
More judges is not a solution
The strongest of the three studies says so itself: several judges together dampen self-preference but do not remove it. Anyone selling a panel as sufficient has not finished reading the study.

What does the study say?

Pombal, Rei and Martins had language models judge each other's work using rubrics — lists of criteria that are individually met or not. Two of those benchmarks, IFEval and LiveCodeBench, have criteria a machine can check: the instruction was followed or it was not, the code runs or it does not. That is exactly where it showed: on criteria the generator fails, a judge wrongly marks them as satisfied more than 50% more often when the text is its own. Several judges together dampen that, but do not remove it.

Yang and colleagues asked the obvious follow-up: does a better model help, then? They compared twenty mainstream models using equal-quality answer pairs, so that bias is separated from the ability to tell a difference. The result is uncomfortable: there is no link between capability and less self-preference, and sometimes it runs the wrong way. What does work is the FORM of the evaluation — a structured, multi-dimensional rubric lowers the bias by 31.5% on average.

And Saraf and colleagues looked at something the other two do not measure: what does the NAME attached to a text do? They had ChatGPT, Gemini and Claude judge each other's texts under four conditions — with no name, with the true name, and with two false names. The label shifts the verdict, even when it is wrong, and not in one direction: one brand raised scores regardless of who actually wrote the text, another lowered them. A false name frequently reversed the preference outright — up to fifty percentage points of difference in the voting and up to twelve in the quality ratings.

What was measuredHowWhat the study does not say
Pombal, Rei & Martins — IFEval, LiveCodeBench and HealthBenchcompared, criterion by criterion, how often a judge marks it as satisfied for its own work versus someone else'snothing about subjective evaluation — the more than 50% comes from machine-verifiable rubrics, and the ten points on HealthBench from subjective ones. Those two figures are not about the same thing and must never be added together
Yang et al. — 20 mainstream language modelspresented equal-quality answer pairs, so that bias stands apart from the ability to tell a differencenothing about any one model — the 31.5% is an average across twenty. And it is a REDUCTION of the bias, not its removal: a rubric makes a judge fairer, not fair
Saraf et al. — ChatGPT, Gemini and Claude, four name conditionsthe same texts judged four times: with no name, with the true name, and with two false onesthat a judge always scores HIGHER once it thinks it knows who wrote the text. The label distorts in BOTH directions, and which way depends on which name is shown. It also covers three models on blog posts; the direction per brand need not be fixed

Pombal, Rei & Martins (arXiv 2604.06996, 8 April 2026) · Yang et al. (arXiv 2604.22891, 24 April 2026) · Saraf et al. (arXiv 2508.21164, 28 August 2025). All three figures were found back on 11 August 2026 in the abstract of the paper itself.

Where can you check this yourself?

  • Self-Preference Bias in Rubric-Based Evaluation of Large Language Models — José Pombal, Ricardo Rei & André F. T. Martins April 8, 2026, revised August 3, 2026

    IFEval and LiveCodeBench — two benchmarks with programmatically verifiable rubrics, plus HealthBench with subjective rubrics

    On rubrics where the generator fails, a judge is more than 50% more likely to incorrectly mark them as satisfied when the output is its own. That holds even when the criteria are entirely objective. Ensembling several judges dampens it without eliminating it. On HealthBench it skews scores by up to 10 points.

    The more than 50% is measured on OBJECTIVE, machine-verifiable rubrics — which makes the finding stronger rather than weaker, since there is nothing to argue about on taste. The 10 points, by contrast, come from a medical benchmark with subjective rubrics; the two figures are therefore not about the same thing and must never be added together.

  • Quantifying and Mitigating Self-Preference Bias of LLM Judges — Jinming Yang, Zheng Hu, Chuxian Qiu, Zhenyu Deng, Xinshan Jiao & Tao Zhou April 24, 2026, revised June 2, 2026

    20 mainstream language models, measured with equal-quality answer pairs so that bias is disentangled from the ability to discriminate

    A more advanced model is not less prone to self-preference — the correlation is absent, and sometimes even negative. What does help is a structured, multi-dimensional evaluation: it reduces the bias by 31.5% on average.

    The 31.5% is an AVERAGE across twenty models and says nothing about any one model. And it is a reduction of the bias, not its removal — a rubric makes a judge fairer, not fair.

  • Quantifying Label-Induced Bias in Large Language Model Self- and Cross-Evaluations — Muskan Saraf, Sajjad Rezvani Boroujeni, Justin Beaudry, Hossein Abedi & Tom Bush August 28, 2025, revised October 9, 2025

    ChatGPT, Gemini and Claude evaluated each other's texts under four conditions: with no name attached, with the true name, and with two false names

    The NAME attached to a text shifts the verdict, even when it is wrong — and not always upwards. The “Claude” label raised scores regardless of who actually wrote the text; the “Gemini” label lowered them. A false name frequently reversed the preference outright: up to 50 percentage points of difference in the voting and up to 12 in the quality ratings. The researchers conclude that blind evaluation and several model families are necessary.

    ⚠ This study does NOT say a judge always scores higher when it thinks it knows who wrote the text — our copy did say that, and the simplification makes the finding weaker. The label distorts in BOTH directions, and which way depends on which name is shown. It also covers three models on blog posts; the direction per brand need not be fixed.

Every link goes to the original publication. What you read above was found back in that text itself on August 11, 2026 — not taken over from someone else's summary.

Why does this matter to you?

Because the copy on your site is increasingly written by an AI, and the check on it usually by the same one. That is where these three studies meet: a model reviewing its own FAQ does not deliver a second reading but a stamp with a number on it. It looks strict, it carries a score, and that is exactly where it leaves standing, more than 50% more often, the mistake it made itself.

And it makes the two usual answers unusable. The first is “take a better model” — that does not help, which is exactly what Yang et al. measure across twenty models. The second is “have several look at it” — that dampens without removing, and it is written in the strongest of the three studies itself. What remains is not a choice of model but a SETUP: judge blind, from a different family, with a rubric instead of an impression.

What does ceeme do with it?

Every proposal we produce goes through a second reading that meets all three requirements. The judge gets no hint of origin — no field names, no system name, no brand — and the instruction states outright that a hunch about origin must not affect the verdict. It comes from a different model family than the writer, because blind only removes the name and not the self-preference. And it scores on 6 axes that each have their own floor, so that a single red axis is enough to reject and a high average cannot paper over a factual error.

And if no judgement can be made, none is given. The second reading can fail — no key, no answer, an answer nothing can be taken from — and the proposal then goes on to the human with the reason attached, instead of with a score nobody calculated. Better no judgement than an invented one: that is the same rule that makes us claim no result without re-measurement.

Frequently asked questions

If an AI does the checking, who checks the AI?

You, and that does not change. The second reading does not replace the signature but comes before it: it holds back what is factually wrong, so that what reaches you is worth reading. A judge approving something publishes nothing by doing so — that is a separate gate, with its own reason attached.

Why cite a study that does not fully rescue your own setup?

Because the limit is the argument. Blind evaluation removes the bias from the name and leaves the self-preference standing; that is what the different model family is for, and it too makes a judge fairer rather than fair. Leaving that out sells a solution where there is a measure — and then the first reader who looks it up is immediately the last one to believe us.

Can I do this myself with the AI I already use?

The setup, yes, and it is described above: have a model from a different family check it, do not tell it who wrote the text, and give it a list of criteria instead of the question whether it is any good. What you do not get is the repetition — the same criteria, every time, with the reason kept per rejected axis, so that across a series you can see whether it is improving. That is what this part exists for, not because it is impossible but because it has to happen again every time.

Read on

Free, no account and no card.