Loupe / Benchmarks / tau²-bench airline

tau²-bench airline

Multi-turn airline customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.

May republish0 incidents recorded
not evaluatedSaturation not assessed80.3 pts · headroom unknown
Text equivalenttau²-bench airline: best reproduced result 80.3 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Best reproduced
0.8029999999999999GPT-5.2xhigh effortIndependentairline

Observed 2025-12-15 · source

Saturation
Remaining headroom has not been measured. Unknown, not zero.
Contamination controls
  • execution against mutable world state
  • deterministic outcome checks rather than an answer key
Licence
MIT Recorded as republishable, so its figures appear on this site.
Update cadence
Manual verified runs; models do not appear automatically
Access path
Aggregate table in the repository README

Incident log

Each entry states what was found, by whom, and when. The kind describes the finding, not anybody’s conduct. An entry without a primary source and a date is not published here.

No incidents recorded in this register.That is a statement about what we have recorded, not a finding that none exist. We have not audited every benchmark, and an absence here is an absence of evidence.