Research-level mathematical problem solving on original, unpublished expert-authored problems.
May republish2 incidents recorded
Text equivalentFrontierMath: best reproduced result 47.2 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Best reproduced
0.4724137931034483Claude Opus 4.8max effortIndependentepoch-evaluated
Remaining headroom has not been measured. Unknown, not zero.
Contamination controls
independently run under a fixed harness
restricted problem content on the FrontierMath tiers
Licence
CC BY 4.0 Recorded as republishable, so its figures appear on this site.
Update cadence
Maintainer-run; models are added when Epoch evaluates them
Access path
CSV inside the Epoch benchmark data archive
Incident log
Each entry states what was found, by whom, and when. The kind describes the finding, not anybody’s conduct. An entry without a primary source and a date is not published here.
2026-06-12Defect audit
The version 2 audit addressed errors in 42 per cent of problems. Scores rose materially, although the ranking between models largely persisted.
Epoch AI disclosed that OpenAI commissioned the 300 problems and owns them, and has access to the problems and their solutions with the exception of a 50-problem holdout. Epoch also said its communication about the arrangement should have been more systematic and transparent.