Benchmark
LM Arena
Human preference between two anonymous model responses to the same prompt, aggregated into an Elo-style rating from crowdsourced pairwise votes.
- Maintainer
- LMArena
- Score
- Elo
- Last checked
- September 2, 2026
Leaderboard
No verified results recorded yet.
Methodology
Visitors submit a prompt, receive two anonymous responses, and vote for the one they prefer. Model identities are revealed only after the vote. Votes are aggregated into a rating with confidence intervals, and the leaderboard moves continuously as new votes arrive.
Important limitations
The rating reflects what voters prefer, not correctness. Style, formatting and length influence outcomes. Vote populations are self-selected and not demographically representative, and a rating is only comparable to others captured on the same date.