Governance register

What each benchmark actually measures.

Every documented case of a leaderboard result not meaning what it appeared to mean was found by reading a policy or comparing an artefact, not by running a detector. So this is a register of checkable facts: what a test measures, how much room is left in it, what stops contamination, who may republish it, and what has gone wrong with it.

Every incident carries a primary source and a date. An entry with neither is not published here. Where a figure has not been measured, the register says so rather than printing a zero.

Ingested, not redistributed

What we hold but may not republish.

These benchmarks inform our own reading and never appear as a figure anywhere on this site. The licence is the reason in every case, and each one links to where that licence was read. Naming them is the point: a register listing only what it may publish would look complete while hiding its own edges.

This list is generated from every restricted record we hold, not maintained by hand. A hand-maintained list of the things you are not showing is the one list nobody notices is wrong.

EQ-Bench Slop Score

Not republishedNot stated

No licence is stated on the results repository, so we have no permission to republish its figures and do not assume one. Nothing about the benchmark is implied: an unstated licence is a fact about paperwork, not about the test.

Licence read at github.com/EQ-bench/eqbench-leaderboard-results

EQ-Bench Judgemark

Not republishedNot stated

No licence is stated on the results repository, so we have no permission to republish its figures and do not assume one. Nothing about the benchmark is implied: an unstated licence is a fact about paperwork, not about the test.

Licence read at github.com/EQ-bench/eqbench-leaderboard-results

L-Eval

Not republishedGPL-3.0

Licensed GPL-3.0. Its reciprocal terms are not ones we are in a position to satisfy for published output, so we hold it for reference and publish no figure from it.

Licence read at github.com/OpenLMLab/LEval

AIME 2025

Not republishedCC BY-NC-SA 4.0

Licensed CC BY-NC-SA 4.0, which does not grant commercial redistribution. We may read it and we may say the set exists; we may not put its numbers on a page. Note this is MathArena, not Epoch AI OTIS Mock AIME, which carries a different licence and is published under its own entry.

Licence read at huggingface.co/datasets/MathArena/aime_2025

MathArena Apex

Not republishedCC BY-NC-SA 4.0

Licensed CC BY-NC-SA 4.0, which does not grant commercial redistribution. We may read it and we may say the set exists; we may not put its numbers on a page. Note this is MathArena, not Epoch AI OTIS Mock AIME, which carries a different licence and is published under its own entry.

Licence read at huggingface.co/datasets/MathArena/aime_2025

Design Arena

Not republishedNo licence granted; automated access prohibited

No licence is granted to users over the platform content, and §17(a) assigns rights to Arcada Labs. Republishing leaderboard data would need express written consent.

Licence read at designarena.ai/terms-and-conditions

The practice record

Facts about benchmarking, not about one benchmark.

These are the governance facts that make the case for a register like this one. Each says what was found, by whom, and when. None of them attributes a motive, because intent is not observable from outside and this product publishes measurements.

2025-04-29

Private pre-release testing on a public leaderboard

A peer-reviewed audit found that undisclosed private testing lets providers try several variants before a public release and withdraw scores they do not like, and estimated a gain of up to 112 per cent relative on ArenaHard from the additional arena data. The finding is about the distributional consequence of a policy, which the audit quantified and the policy itself did not.

Singh et al., NeurIPS 2025 (peer-reviewed)

2025-05-01

The operator’s reply, published beside the audit

LMArena responded that the policy had been publicly available since March 2024, that any provider may submit as many public and private variants as capacity allows, and that larger labs submit more models because they build more. It also announced changes: clearer marking of retired models, and provisional scoring until fresh post-release votes accumulate where many models underwent parallel pre-release testing.

LMArena (first-party)

2025-04-06

The evaluated system was not the released system

Meta submitted a variant described as experimental and tuned for human preference, which ranked second, rather than the weights it released for download. Meta has said it described the entry as an experimental chat version. The checkable claim is the artefact difference between what was evaluated and what shipped, which nobody has disputed.

TechCrunch (reporting)

2026-05-12

A benchmark that could not be built

MathArena reported that it could not ship a 2026 set at all, and said the earlier set could by then be contaminated. The maintainers withdrew a planned benchmark on their own assessment of its problem supply. That is task-source exhaustion reported by the people building the tasks, and no detector was involved in establishing it.

MathArena (first-party)