Benchmark

MMLU-Pro

Multiple-choice accuracy across fourteen academic and professional domains, using ten answer options per question rather than the four used by the original MMLU.

Maintainer
TIGER-Lab
Score
%
Last checked
September 2, 2026

Leaderboard

No verified results recorded yet.

Methodology

The dataset extends MMLU with harder, more reasoning-intensive questions and expands each question to ten options, which lowers the score obtainable by guessing. Noisy and trivially answerable items from the original set were filtered out during construction.

Important limitations

It remains a multiple-choice test of recall and elimination rather than of open-ended reasoning or real task performance. Prompt format and whether chain-of-thought is permitted change results by several points, so harness details matter when comparing.