Governance register

What each benchmark actually measures.

Every documented case of a leaderboard result not meaning what it appeared to mean was found by reading a policy or comparing an artefact, not by running a detector. So this is a register of checkable facts: what a test measures, how much room is left in it, what stops contamination, who may republish it, and what has gone wrong with it.

Every incident carries a primary source and a date. An entry with neither is not published here. Where a figure has not been measured, the register says so rather than printing a zero.

Nothing recorded under this filter.That is a statement about the register, not about the benchmarks: nothing has been assessed this way, rather than everything having been ruled out.
The practice record

Facts about benchmarking, not about one benchmark.

These are the governance facts that make the case for a register like this one. Each says what was found, by whom, and when. None of them attributes a motive, because intent is not observable from outside and this product publishes measurements.

2025-04-29

Private pre-release testing on a public leaderboard

A peer-reviewed audit found that undisclosed private testing lets providers try several variants before a public release and withdraw scores they do not like, and estimated a gain of up to 112 per cent relative on ArenaHard from the additional arena data. The finding is about the distributional consequence of a policy, which the audit quantified and the policy itself did not.

Singh et al., NeurIPS 2025 (peer-reviewed)

2025-05-01

The operator’s reply, published beside the audit

LMArena responded that the policy had been publicly available since March 2024, that any provider may submit as many public and private variants as capacity allows, and that larger labs submit more models because they build more. It also announced changes: clearer marking of retired models, and provisional scoring until fresh post-release votes accumulate where many models underwent parallel pre-release testing.

LMArena (first-party)

2025-04-06

The evaluated system was not the released system

Meta submitted a variant described as experimental and tuned for human preference, which ranked second, rather than the weights it released for download. Meta has said it described the entry as an experimental chat version. The checkable claim is the artefact difference between what was evaluated and what shipped, which nobody has disputed.

TechCrunch (reporting)

2026-05-12

A benchmark that could not be built

MathArena reported that it could not ship a 2026 set at all, and said the earlier set could by then be contaminated. The maintainers withdrew a planned benchmark on their own assessment of its problem supply. That is task-source exhaustion reported by the people building the tasks, and no detector was involved in establishing it.

MathArena (first-party)