Benchmark
LiveBench
Aggregate accuracy across reasoning, coding, mathematics, data analysis, language and instruction-following tasks that are refreshed regularly to limit contamination.
- Maintainer
- Abacus.AI, NYU and collaborators
- Score
- %
- Last checked
- September 2, 2026
Leaderboard
No verified results recorded yet.
Methodology
Questions are drawn from recent sources such as newly published papers, competitions and news, and are replaced on a rolling schedule. Every task has an objective, automatically verifiable ground-truth answer, so scoring requires no model judge and no human rater.
Important limitations
Because the question set changes over time, scores from different refresh cycles are not directly comparable. The aggregate figure hides large variation between categories, and a single headline number can conceal weakness in a specific skill.