Loupe / Benchmarks / SWE-bench Verified

SWE-bench Verified

Resolving real Python repository issues by generating patches that must pass repository tests.

May republish2 incidents recorded
not evaluatedSaturation not assessed83.5 pts · headroom unknown
Text equivalentSWE-bench Verified: best reproduced result 83.5 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Best reproduced
0.8347107438016529Claude Opus 4.7max effortIndependentepoch-evaluated

Observed 2026-04-20 · source

Saturation
Remaining headroom has not been measured. Unknown, not zero.
Contamination controls
  • independently run under a fixed harness
  • restricted problem content on the FrontierMath tiers
Licence
CC BY 4.0 Recorded as republishable, so its figures appear on this site.
Update cadence
Maintainer-run; models are added when Epoch evaluates them
Access path
CSV inside the Epoch benchmark data archive

Incident log

Each entry states what was found, by whom, and when. The kind describes the finding, not anybody’s conduct. An entry without a primary source and a date is not published here.