Loupe / Benchmarks / LMArena Text Arena

LMArena Text Arena

Crowd preference over blind pairwise responses to open-ended text prompts, reported as an Elo-style rating with an interval.

May republish0 incidents recorded

No fixed ceiling

Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it.

Best reproduced
1496.6257056431264Claude Opus 4.6high effortIndependenttext/overall

Observed 2026-07-30 · source

Saturation
Remaining headroom has not been measured. Unknown, not zero.
Contamination controls
  • fresh user-generated prompts
  • anonymous model identities before voting
  • no reusable static test set
Licence
CC BY 4.0 Recorded as republishable, so its figures appear on this site.
Update cadence
Continuous, snapshotted
Access path
Hugging Face datasets-server JSON

Incident log

Each entry states what was found, by whom, and when. The kind describes the finding, not anybody’s conduct. An entry without a primary source and a date is not published here.

No incidents recorded in this register.That is a statement about what we have recorded, not a finding that none exist. We have not audited every benchmark, and an absence here is an absence of evidence.