Agents completing realistic command-line tasks in containerised environments, checked by deterministic tests.
May republish2 incidents recorded
Text equivalentTerminal-Bench 2.0: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Best reported
No result recorded in this register.
Saturation
Remaining headroom has not been measured. Unknown, not zero.
Contamination controls
canary strings
checks that tests and oracle solutions are absent from agent-visible containers
audited configurations and verified trajectories
Licence
Apache-2.0 Recorded as republishable, so its figures appear on this site.
Update cadence
Submission-driven; entries are audited rather than published automatically
Access path
Rendered leaderboard; submission trajectories on Hugging Face
Incident log
Each entry states what was found, by whom, and when. The kind describes the finding, not anybody’s conduct. An entry without a primary source and a date is not published here.
2026-04-08Defect audit
An audit corrected issues in 28 of the 89 tasks, changing absolute scores by up to 12 points.
The same audit demonstrated that a fake curl wrapper could score 100 per cent without completing any task, which measures what the harness checks rather than what any model can do.