Benchmark

LM Arena

Human preference between two anonymous model responses to the same prompt, aggregated into an Elo-style rating from crowdsourced pairwise votes.

Maintainer
LMArena
Score
Elo
Last checked
September 2, 2026

Leaderboard

No verified results recorded yet.

Methodology

Visitors submit a prompt, receive two anonymous responses, and vote for the one they prefer. Model identities are revealed only after the vote. Votes are aggregated into a rating with confidence intervals, and the leaderboard moves continuously as new votes arrive.

Important limitations

The rating reflects what voters prefer, not correctness. Style, formatting and length influence outcomes. Vote populations are self-selected and not demographically representative, and a rating is only comparable to others captured on the same date.