Governance register

What each benchmark actually measures.

Every documented case of a leaderboard result not meaning what it appeared to mean was found by reading a policy or comparing an artefact, not by running a detector. So this is a register of checkable facts: what a test measures, how much room is left in it, what stops contamination, who may republish it, and what has gone wrong with it.

Every incident carries a primary source and a date. An entry with neither is not published here. Where a figure has not been measured, the register says so rather than printing a zero.

LMArena Text Arena

May republish

Crowd preference over blind pairwise responses to open-ended text prompts, reported as an Elo-style rating with an interval.

CC BY 4.0Continuous, snapshotted0 incidents recorded

No fixed ceiling

Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.

Human preference between anonymously generated front-end websites.

CC BY 4.0Continuous, snapshotted0 incidents recorded

No fixed ceiling

Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.

Human preference over responses to image-bearing prompts.

CC BY 4.0Continuous, snapshotted0 incidents recorded

No fixed ceiling

Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.

Aider Polyglot

May republish

Instruction-following code editing across 225 Exercism exercises in six languages, scored on pass rate after one retry and on whether the edit was well formed.

Apache-2.0Maintainer-run; new models do not appear automatically0 incidents recorded
not evaluatedSaturation not assessed76.9 pts · headroom unknown
Text equivalentAider Polyglot: best reproduced result 76.9 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

tau²-bench airline

May republish

Multi-turn airline customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.

MITManual verified runs; models do not appear automatically0 incidents recorded
not evaluatedSaturation not assessed80.3 pts · headroom unknown
Text equivalenttau²-bench airline: best reproduced result 80.3 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

tau²-bench retail

May republish

Multi-turn retail customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.

MITManual verified runs; models do not appear automatically0 incidents recorded
not evaluatedSaturation not assessed79.7 pts · headroom unknown
Text equivalenttau²-bench retail: best reproduced result 79.7 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

tau²-bench telecom

May republish

Multi-turn telecom customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.

MITManual verified runs; models do not appear automatically0 incidents recorded
not evaluatedSaturation not assessed89.5 pts · headroom unknown
Text equivalenttau²-bench telecom: best reproduced result 89.5 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

GPQA Diamond

May republish

Expert-level multiple-choice reasoning in physics, chemistry and biology.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessed93.9 pts · headroom unknown
Text equivalentGPQA Diamond: best reproduced result 93.9 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

FrontierMath

May republish

Research-level mathematical problem solving on original, unpublished expert-authored problems.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them2 incidents recorded
not evaluatedSaturation not assessed47.2 pts · headroom unknown
Text equivalentFrontierMath: best reproduced result 47.2 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

FrontierMath Tier 4

May republish

The hardest FrontierMath tier, using restricted materials rather than a public answer key.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessed31.3 pts · headroom unknown
Text equivalentFrontierMath Tier 4: best reproduced result 31.3 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

MATH Level 5

May republish

The hardest tier of the MATH competition-mathematics dataset.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentMATH Level 5: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SWE-bench Verified

May republish

Resolving real Python repository issues by generating patches that must pass repository tests.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them2 incidents recorded
not evaluatedSaturation not assessed83.5 pts · headroom unknown
Text equivalentSWE-bench Verified: best reproduced result 83.5 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SimpleQA Verified

May republish

Short-answer factual accuracy and hallucination resistance on fact-seeking questions.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessed71.6 pts · headroom unknown
Text equivalentSimpleQA Verified: best reproduced result 71.6 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

Fiction.LiveBench

May republish

Recall and comprehension over long serialized fiction.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentFiction.LiveBench: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SimpleBench

May republish

Everyday commonsense, spatial and social reasoning with adversarial linguistic traps.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentSimpleBench: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

Instruction-following code editing, as collected by Epoch from external reports.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentAider Polyglot (externally reported): no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SWE-rebench

May republish

Coding-agent issue resolution on continuously collected recent GitHub issue/PR pairs under a fixed scaffold, repeated five times.

CC BY 4.0Monthly rolling splits1 incident recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentSWE-rebench: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

NIAH v2

May republish

Long-context retrieval across context length and needle depth, including multi-fact and UUID-chain tasks.

MIT0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentNIAH v2: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

GraphWalks

May republish

Multi-hop reasoning over graphs represented as edge lists in long context.

MIT0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentGraphWalks: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

RULER

May republish

The effective context size of long-context models via multi-hop tracing and aggregation beyond simple retrieval.

Apache-2.00 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentRULER: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

IFEval-FC

May republish

Strict instruction and format adherence, scored by regex and JSON-schema validation rather than a judge.

Apache-2.00 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentIFEval-FC: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SafetyBench

May republish

Safety understanding across seven categories, in Chinese and English.

MIT0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentSafetyBench: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

HaluEval

May republish

Whether a model recognises and avoids hallucinated content, against human-annotated samples.

MIT0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentHaluEval: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

Terminal-Bench 2.0

May republish

Agents completing realistic command-line tasks in containerised environments, checked by deterministic tests.

Apache-2.0Submission-driven; entries are audited rather than published automatically2 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentTerminal-Bench 2.0: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

BrowseComp

May republish

Persistent agentic web research on questions whose answers are easy to verify and hard to locate.

MIT (evaluation tooling)No unified cadence; scores appear when vendors or independent evaluators publish runs1 incident recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentBrowseComp: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

DeepSWE

May republish

Whether a coding agent can complete original, long-horizon engineering tasks in real repositories, verified by hand-written tests that check behaviour rather than implementation details.

Apache-2.00 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentDeepSWE: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Ingested, not redistributed

What we hold but may not republish.

These benchmarks inform our own reading and never appear as a figure anywhere on this site. The licence is the reason in every case, and each one links to where that licence was read. Naming them is the point: a register listing only what it may publish would look complete while hiding its own edges.

This list is generated from every restricted record we hold, not maintained by hand. A hand-maintained list of the things you are not showing is the one list nobody notices is wrong.

EQ-Bench Slop Score

Not republishedNot stated

No licence is stated on the results repository, so we have no permission to republish its figures and do not assume one. Nothing about the benchmark is implied: an unstated licence is a fact about paperwork, not about the test.

Licence read at github.com/EQ-bench/eqbench-leaderboard-results

EQ-Bench Judgemark

Not republishedNot stated

No licence is stated on the results repository, so we have no permission to republish its figures and do not assume one. Nothing about the benchmark is implied: an unstated licence is a fact about paperwork, not about the test.

Licence read at github.com/EQ-bench/eqbench-leaderboard-results

L-Eval

Not republishedGPL-3.0

Licensed GPL-3.0. Its reciprocal terms are not ones we are in a position to satisfy for published output, so we hold it for reference and publish no figure from it.

Licence read at github.com/OpenLMLab/LEval

AIME 2025

Not republishedCC BY-NC-SA 4.0

Licensed CC BY-NC-SA 4.0, which does not grant commercial redistribution. We may read it and we may say the set exists; we may not put its numbers on a page. Note this is MathArena, not Epoch AI OTIS Mock AIME, which carries a different licence and is published under its own entry.

Licence read at huggingface.co/datasets/MathArena/aime_2025

MathArena Apex

Not republishedCC BY-NC-SA 4.0

Licensed CC BY-NC-SA 4.0, which does not grant commercial redistribution. We may read it and we may say the set exists; we may not put its numbers on a page. Note this is MathArena, not Epoch AI OTIS Mock AIME, which carries a different licence and is published under its own entry.

Licence read at huggingface.co/datasets/MathArena/aime_2025

Design Arena

Not republishedNo licence granted; automated access prohibited

No licence is granted to users over the platform content, and §17(a) assigns rights to Arcada Labs. Republishing leaderboard data would need express written consent.

Licence read at designarena.ai/terms-and-conditions

Evaluated, not ingested

What we looked at and refused.

The section above lists sources we hold and may not republish. These we never took at all. Each was evaluated against a specific need, its licence read at a link you can follow, and rejected on a stated ground. Recording the refusals is what stops the next person repeating the search and reaching the same wall.

Every entry carries the date it was checked, because a refusal is a claim about someone else’s work at a moment in time and those change. None of these cards carries a figure: naming a source we may not republish is fine, printing its measurement is not.

No redistribution grant

Vercel AI Gateway

No redistribution grantUnspecified

Publishes a genuine token rate per provider endpoint, on a rolling window, alongside a latency figure. It is the closest thing to the measurement this product wanted. No published term grants redistribution, and this repo already carries the source as restricted on that basis. Two further problems stand even if the licence were resolved: the rate is measured per provider endpoint, so one model has several and the field it would fill holds a single value, and the measurement states no effort.

What would change this: A published licence granting redistribution, plus a stated effort. The per-endpoint granularity would still need a documented rule for choosing among a model’s endpoints.

Checked 2026-07-31vercel.com/docs/ai-gateway

Artificial Analysis

No redistribution grantInternal use only; no redistribution (free tier)

Publishes a median time-to-first-token and an output rate across hosted models, which is the right quantity for the right population. The free tier is marked internal use only with no redistribution, and its endpoint refuses an unauthenticated request. Redistribution is available on the paid Commercial tier, which makes this a commercial decision rather than an engineering one.

What would change this: A Commercial licence. That is a purchase, not a code change, and it would also need the effort question settled before the figures could be compared across models.

Checked 2026-07-31artificialanalysis.ai/documentation

Publishes no rate

models.dev

Publishes no rateMITLicence permits reuse

A permissively licensed catalogue this product already reads for other fields. Walking its published payload for any rate-bearing key returns nothing: the only matches are model names containing the word throughput, not fields carrying a figure.

What would change this: A throughput field in the published catalogue. The licence already permits it, so this is the cheapest of the set to adopt if it ever appears.

Checked 2026-07-31github.com/sst/models.dev

LiteLLM

Publishes no rateMITLicence permits reuse

A permissively licensed price map covering the hosted models this product ranks. It carries pricing and context limits and no performance measurement of any kind.

What would change this: A throughput field in the price map. Nothing about the licence stands in the way.

Checked 2026-07-31github.com/BerriAI/litellm/blob/main/LICENSE

OpenRouter

Publishes no rateUnspecified

Its public model listing was walked for any rate-bearing key and carries none. Recorded because a router serving many providers is the obvious place to look for a rate, so the next person will look here too. The licence question was never reached.

What would change this: A published rate in the model listing, and then a redistribution term, in that order.

Checked 2026-07-31openrouter.ai/docs

Helicone

Publishes no rateUnspecified

Its public model registry was walked for any rate-bearing key and carries none. Recorded for the same reason as the router above: an observability vendor looks like it should publish a latency figure, and it does not.

What would change this: A published rate in the public registry, and then a redistribution term.

Checked 2026-07-31docs.helicone.ai

Wrong population

llm-speed

Wrong populationCC BY 4.0 (data); Apache-2.0 (code)Licence permits reuse

The only source in this search that both publishes a decode rate and grants redistribution: the data is Creative Commons Attribution, explicitly permitting commercial reuse with a credit and a link. It was refused on coverage, not licensing. Every run in the published corpus is local inference — the backends are llama.cpp, mlx and ollama, and the accelerators are consumer GPUs and Apple Silicon. It carries no figure for any hosted model this product ranks, and its row schema has no field in which a hosted endpoint could be recorded. Re-measured on the date shown, against both the bulk corpus and the live API, with the parse proved against the local rows it does contain.

What would change this: Hosted-API rows with a stated effort. This is the strongest candidate found and the decision should be revisited if that appears. The unit would still need a rule: the corpus is keyed by hardware, so one model carries as many rates as there are machines.

Checked 2026-07-31llm-speed.com/data

The practice record

Facts about benchmarking, not about one benchmark.

These are the governance facts that make the case for a register like this one. Each says what was found, by whom, and when. None of them attributes a motive, because intent is not observable from outside and this product publishes measurements.

2025-04-29

Private pre-release testing on a public leaderboard

A peer-reviewed audit found that undisclosed private testing lets providers try several variants before a public release and withdraw scores they do not like, and estimated a gain of up to 112 per cent relative on ArenaHard from the additional arena data. The finding is about the distributional consequence of a policy, which the audit quantified and the policy itself did not.

Singh et al., NeurIPS 2025 (peer-reviewed)

2025-05-01

The operator’s reply, published beside the audit

LMArena responded that the policy had been publicly available since March 2024, that any provider may submit as many public and private variants as capacity allows, and that larger labs submit more models because they build more. It also announced changes: clearer marking of retired models, and provisional scoring until fresh post-release votes accumulate where many models underwent parallel pre-release testing.

LMArena (first-party)

2025-04-06

The evaluated system was not the released system

Meta submitted a variant described as experimental and tuned for human preference, which ranked second, rather than the weights it released for download. Meta has said it described the entry as an experimental chat version. The checkable claim is the artefact difference between what was evaluated and what shipped, which nobody has disputed.

TechCrunch (reporting)

2026-05-12

A benchmark that could not be built

MathArena reported that it could not ship a 2026 set at all, and said the earlier set could by then be contaminated. The maintainers withdrew a planned benchmark on their own assessment of its problem supply. That is task-source exhaustion reported by the people building the tasks, and no detector was involved in establishing it.

MathArena (first-party)