Your LLM Judge is lying to you: the measurement crisis in AI benchmarks

— LLM, Benchmarks, Evaluation, Python

A new preregistered study shows that LLM judges — the instruments behind every modern benchmark — are fundamentally unreliable. Same prompt, same model, different answer. Here's what it means and how to think about it.

The instrument you trust is broken

Every major LLM leaderboard today relies on the same trick: instead of hiring thousands of human annotators, you ask another LLM to judge which output is better. GPT-4 judges AlpacaEval. Claude judges Arena. The approach is fast, cheap, and scales effortlessly. It has become the invisible backbone of how we measure progress in AI.

But there is a deeply uncomfortable assumption buried in this workflow: the same request, sent to the same model name, produces the same judgment tomorrow. If that assumption fails – if the instrument drifts – then every benchmark score built on it is noise dressed up as signal.

A new preregistered study by Zhu and Zhang (Zhu and Zhang 2026) sets out to test exactly this assumption. The results are stark: the assumption fails, badly.

The experiment

The authors ran two preregistered campaigns – meaning every threshold, metric, and decision rule was locked in advance, before any data was collected. This is the gold standard for avoiding cherry-picking. Across 52,988 audited request attempts to black-box LLM endpoints, they measured two things:

  1. Same-window repeatability: Send the same ranking prompt twice in quick succession. How often do you get the same ranking?
  2. Next-day reproducibility: Send the exact same bytes tomorrow. How often does the ranking match?

Both were measured with Spearman rank correlation – where 1.0 means perfect agreement and 0.0 means no relationship at all.

Required vs achieved agreement for LLM judges. Neither metric passes its own quality gate.
Required vs achieved agreement for LLM judges. Neither metric passes its own quality gate.

The numbers are unforgiving. Same-window repeatability came in at 0.40 – a coin flip has more consistency than two identical calls to the same judge. Next-day reproducibility reached 0.78, better but still well below the 0.99 the authors required. And these are not fluke results: the execution record (response times, token counts) was at ceiling, meaning the infrastructure was working fine. The instability is in the model itself.

Three mechanisms behind the noise

The paper identifies three distinct failure modes, each worth understanding:

Label-to-meaning mapping bias

When you ask an LLM to rate a response on a 1–5 scale, the model does not apply a fixed psychometric function. The meaning of “4 out of 5” drifts between calls, between contexts, and between days. The label is not a calibrated instrument – it is a sampled guess from a distribution that shifts with every request.

Signal below the noise floor

Many benchmark comparisons involve outputs that differ by tiny margins – a few percentage points on HumanEval, a fraction of a MMLU category. The study found that these candidate gaps sit seven orders of magnitude below the judge’s own noise floor. It is like trying to weigh a feather on a bathroom scale that fluctuates by kilograms.

Non-reproducible rankings

Even byte-identical inputs – character-for-character identical prompts sent to the same endpoint – returned different rankings. This is not a temperature issue or a sampling artifact. The serving infrastructure itself introduces non-determinism that exact-permutation readouts compound.

The three failure mechanisms and their relative impact on benchmark reliability.
The three failure mechanisms and their relative impact on benchmark reliability.

Providers are equally unreliable

A natural response is: “Just use a better provider.” The study tested four major LLM API providers and found that all of them share the same floor. Median Spearman agreement ranged from 0.74 to 0.88 across providers – none approaching the 0.99 threshold. And critically, none of the metadata fields the providers expose (model version, region, endpoint type) predicted which requests would fail.

Median Spearman agreement across four LLM API providers. All fall short of the 0.99 reliability threshold.
Median Spearman agreement across four LLM API providers. All fall short of the 0.99 reliability threshold.

What this means for benchmarks

This is not an academic curiosity. It has immediate, concrete consequences for every leaderboard that uses LLM-as-judge evaluation:

The paper distills its findings into a three-level snapshot-identity ladder: a framework for specifying what level of reproducibility your evaluation actually needs, and eight design rules for getting there.

A practical takeaway: measure your instrument

The single most actionable recommendation from the paper is almost embarrassingly simple: measure your instrument before you freeze any gate on it. A pilot study at roughly 2% of the study’s call volume would have caught both unreachable quality gates in advance.

Here is a minimal Python sketch of what that looks like – send the same prompt twice and check the agreement:

Simulated judge reproducibility: distribution of Spearman rho across 500 paired evaluations. The median sits around 0.75, far below the 0.99 threshold needed for reliable rankings.
Simulated judge reproducibility: distribution of Spearman rho across 500 paired evaluations. The median sits around 0.75, far below the 0.99 threshold needed for reliable rankings.

The simulation is simple, but it captures the core finding: when the judge introduces per-evaluation noise of sigma = 1.2 on a 1–10 scale, the resulting Spearman correlations cluster around 0.75 – nowhere near the 0.99 you need to trust fine-grained rankings.

The broader lesson

The LLM-as-judge paradigm is not broken beyond repair. It is broken beyond trust unless you actively verify it. The paper’s contribution is not that LLMs are bad judges – it is that the community has been treating a noisy, drifting instrument as if it were a calibrated voltmeter.

For benchmark creators, the path forward is clear:

  1. Pre-register your evaluation protocol – fix thresholds before seeing results
  2. Measure instrument reliability – run paired evaluations and report agreement
  3. Report confidence intervals, not point estimates – a leaderboard with error bars is more honest than one without
  4. Consider ensemble judges – multiple independent judges can average out individual noise

The age of “GPT-4 said it’s better, so it is” is over. The instrument needs to be validated before we trust the measurement.


Paper: Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints by Haoyaun Zhu and Jie Zhang (2026).

Zhu, Haoyaun, and Jie Zhang. 2026. Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints. https://arxiv.org/abs/2609.04198v1.