Loupe / Benchmarks / SWE-bench Verified
Resolving real Python repository issues by generating patches that must pass repository tests.
Observed 2026-04-20 · source
Each entry states what was found, by whom, and when. The kind describes the finding, not anybody’s conduct. An entry without a primary source and a date is not published here.
OpenAI stopped evaluating on SWE-bench Verified. It audited 138 hard tasks and reported that 59.4 per cent carried material test or prompt defects, and that every frontier model it tested could reproduce some gold patch or prompt detail.
openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
A forensic probe reported recovering file paths from issue text alone at up to 76 per cent on this set, against up to 53 per cent elsewhere, and reproducing 5-gram functions at 35 per cent against 18 per cent. The probe measures recovery; it does not establish how the material was learned.