Benchmark

LiveBench

Aggregate accuracy across reasoning, coding, mathematics, data analysis, language and instruction-following tasks that are refreshed regularly to limit contamination.

Maintainer
Abacus.AI, NYU and collaborators
Score
%
Last checked
September 2, 2026

Leaderboard

No verified results recorded yet.

Methodology

Questions are drawn from recent sources such as newly published papers, competitions and news, and are replaced on a rolling schedule. Every task has an objective, automatically verifiable ground-truth answer, so scoring requires no model judge and no human rater.

Important limitations

Because the question set changes over time, scores from different refresh cycles are not directly comparable. The aggregate figure hides large variation between categories, and a single headline number can conceal weakness in a specific skill.