Governance register

What each benchmark actually measures.

Every documented case of a leaderboard result not meaning what it appeared to mean was found by reading a policy or comparing an artefact, not by running a detector. So this is a register of checkable facts: what a test measures, how much room is left in it, what stops contamination, who may republish it, and what has gone wrong with it.

Every incident carries a primary source and a date. An entry with neither is not published here. Where a figure has not been measured, the register says so rather than printing a zero.

Evaluated, not ingested

What we looked at and refused.

The section above lists sources we hold and may not republish. These we never took at all. Each was evaluated against a specific need, its licence read at a link you can follow, and rejected on a stated ground. Recording the refusals is what stops the next person repeating the search and reaching the same wall.

Every entry carries the date it was checked, because a refusal is a claim about someone else’s work at a moment in time and those change. None of these cards carries a figure: naming a source we may not republish is fine, printing its measurement is not.

No redistribution grant

Vercel AI Gateway

No redistribution grantUnspecified

Publishes a genuine token rate per provider endpoint, on a rolling window, alongside a latency figure. It is the closest thing to the measurement this product wanted. No published term grants redistribution, and this repo already carries the source as restricted on that basis. Two further problems stand even if the licence were resolved: the rate is measured per provider endpoint, so one model has several and the field it would fill holds a single value, and the measurement states no effort.

What would change this: A published licence granting redistribution, plus a stated effort. The per-endpoint granularity would still need a documented rule for choosing among a model’s endpoints.

Checked 2026-07-31vercel.com/docs/ai-gateway

Artificial Analysis

No redistribution grantInternal use only; no redistribution (free tier)

Publishes a median time-to-first-token and an output rate across hosted models, which is the right quantity for the right population. The free tier is marked internal use only with no redistribution, and its endpoint refuses an unauthenticated request. Redistribution is available on the paid Commercial tier, which makes this a commercial decision rather than an engineering one.

What would change this: A Commercial licence. That is a purchase, not a code change, and it would also need the effort question settled before the figures could be compared across models.

Checked 2026-07-31artificialanalysis.ai/documentation

Publishes no rate

models.dev

Publishes no rateMITLicence permits reuse

A permissively licensed catalogue this product already reads for other fields. Walking its published payload for any rate-bearing key returns nothing: the only matches are model names containing the word throughput, not fields carrying a figure.

What would change this: A throughput field in the published catalogue. The licence already permits it, so this is the cheapest of the set to adopt if it ever appears.

Checked 2026-07-31github.com/sst/models.dev

LiteLLM

Publishes no rateMITLicence permits reuse

A permissively licensed price map covering the hosted models this product ranks. It carries pricing and context limits and no performance measurement of any kind.

What would change this: A throughput field in the price map. Nothing about the licence stands in the way.

Checked 2026-07-31github.com/BerriAI/litellm/blob/main/LICENSE

OpenRouter

Publishes no rateUnspecified

Its public model listing was walked for any rate-bearing key and carries none. Recorded because a router serving many providers is the obvious place to look for a rate, so the next person will look here too. The licence question was never reached.

What would change this: A published rate in the model listing, and then a redistribution term, in that order.

Checked 2026-07-31openrouter.ai/docs

Helicone

Publishes no rateUnspecified

Its public model registry was walked for any rate-bearing key and carries none. Recorded for the same reason as the router above: an observability vendor looks like it should publish a latency figure, and it does not.

What would change this: A published rate in the public registry, and then a redistribution term.

Checked 2026-07-31docs.helicone.ai

Wrong population

llm-speed

Wrong populationCC BY 4.0 (data); Apache-2.0 (code)Licence permits reuse

The only source in this search that both publishes a decode rate and grants redistribution: the data is Creative Commons Attribution, explicitly permitting commercial reuse with a credit and a link. It was refused on coverage, not licensing. Every run in the published corpus is local inference — the backends are llama.cpp, mlx and ollama, and the accelerators are consumer GPUs and Apple Silicon. It carries no figure for any hosted model this product ranks, and its row schema has no field in which a hosted endpoint could be recorded. Re-measured on the date shown, against both the bulk corpus and the live API, with the parse proved against the local rows it does contain.

What would change this: Hosted-API rows with a stated effort. This is the strongest candidate found and the decision should be revisited if that appears. The unit would still need a rule: the corpus is keyed by hardware, so one model carries as many rates as there are machines.

Checked 2026-07-31llm-speed.com/data

The practice record

Facts about benchmarking, not about one benchmark.

These are the governance facts that make the case for a register like this one. Each says what was found, by whom, and when. None of them attributes a motive, because intent is not observable from outside and this product publishes measurements.

2025-04-29

Private pre-release testing on a public leaderboard

A peer-reviewed audit found that undisclosed private testing lets providers try several variants before a public release and withdraw scores they do not like, and estimated a gain of up to 112 per cent relative on ArenaHard from the additional arena data. The finding is about the distributional consequence of a policy, which the audit quantified and the policy itself did not.

Singh et al., NeurIPS 2025 (peer-reviewed)

2025-05-01

The operator’s reply, published beside the audit

LMArena responded that the policy had been publicly available since March 2024, that any provider may submit as many public and private variants as capacity allows, and that larger labs submit more models because they build more. It also announced changes: clearer marking of retired models, and provisional scoring until fresh post-release votes accumulate where many models underwent parallel pre-release testing.

LMArena (first-party)

2025-04-06

The evaluated system was not the released system

Meta submitted a variant described as experimental and tuned for human preference, which ranked second, rather than the weights it released for download. Meta has said it described the entry as an experimental chat version. The checkable claim is the artefact difference between what was evaluated and what shipped, which nobody has disputed.

TechCrunch (reporting)

2026-05-12

A benchmark that could not be built

MathArena reported that it could not ship a 2026 set at all, and said the earlier set could by then be contaminated. The maintainers withdrew a planned benchmark on their own assessment of its problem supply. That is task-source exhaustion reported by the people building the tasks, and no detector was involved in establishing it.

MathArena (first-party)