Loupe / Benchmarks / DeepSWE

DeepSWE

Whether a coding agent can complete original, long-horizon engineering tasks in real repositories, verified by hand-written tests that check behaviour rather than implementation details.

May republish0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentDeepSWE: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Best reported
No result recorded in this register.
Saturation
Remaining headroom has not been measured. Unknown, not zero.
Contamination controls
No controls are recorded in this register.
Licence
Apache-2.0 Recorded as republishable, so its figures appear on this site.
Update cadence
Not recorded.
Access path
HTML leaderboard; tasks and verifiers on GitHub

Incident log

Each entry states what was found, by whom, and when. The kind describes the finding, not anybody’s conduct. An entry without a primary source and a date is not published here.

No incidents recorded in this register.That is a statement about what we have recorded, not a finding that none exist. We have not audited every benchmark, and an absence here is an absence of evidence.