Two results, one set of models

Mercor published human baselines on Thursday for a set of simplified accounting tasks, and the comparison is lopsided. Twelve junior accountants, all licensed CPAs with an average of about five and a half years of experience, averaged 37 per cent across four month-end close scenarios, with individual results ranging from zero to roughly 90 per cent and taking between 30 and 180 minutes per task. Mercor says Claude Opus 5 scored 100 per cent on all twenty of its attempts, finishing each in under ten minutes.

Eighteen months earlier, Mercor says, models scored below the human average on the same tasks.

That is one number. The other is on the full benchmark these tasks were simplified from, and it is considerably less flattering.

What the harder version says

APEX-Accounting, built by Mercor with the finance software company Ramp, is 160 tasks set across ten simulated companies, with an average of 7.5 input files per task and 2,186 grading criteria written by accounting professionals. On that suite, Mercor’s leaderboard puts Claude Opus 5.5 at 61.8 per cent, Fable 5.1 at 61.0 per cent and GPT-6 Astra at 57.9 per cent, each with an error bar near four points.

Rows of binders on office shelves
APEX-Accounting sets 160 tasks across ten simulated companies. Illustrative photo. Zulfugar Karimov · pexels · Pexels License

The consistency figures are the harder read. Every model is run eight times per task. The most consistent model got a task right on all eight runs just 2.6 per cent of the time, and no model fully solved close to 60 per cent of the tasks on any run. Mercor says roughly seven in ten failures by the top three models are reasoning failures rather than failures to find a file or follow an instruction.

Why the gap between the two numbers is the story

The simplified tasks isolate what current models are unambiguously good at: reading a lot of files carefully, not losing a number, and doing what the instruction said. Those are real accounting skills, and a human doing them for the eleventh hour in a row does them worse than a machine that has no eleventh hour.

What the simplified tasks remove is everything that makes a close hard — the judgement about which number is wrong, the call to the person who booked it, the knowledge that this particular client always posts the accrual late. Mercor says so explicitly: the authors note the tasks measure only a subset of accounting skills, and exclude client communication and tacit context.

Two people talking across a cafe table with coffee cups
The benchmark excludes client communication and tacit context. Illustrative photo. August de Richelieu · pexels · Pexels License

That caveat is doing a great deal of work, and it is worth keeping attached to the 100 per cent.

Whose benchmark this is

Neither number is independent. Mercor is a marketplace that sells expert human labour to AI labs for exactly this kind of data, and Ramp sells finance software. A benchmark showing that models are close but not there is a benchmark that argues for more expert data, which is Mercor’s product. That does not make the measurements wrong — the grading rubrics were written by professionals from large accounting firms, and an LM judge that the company says matches human graders 97 per cent of the time does the scoring — but it is the vendor’s own evaluation of a market it sells into.

The tasks are also fixed and public, which means their value decays as they enter training data. A score of 61.8 per cent in October 2026 is a statement about this month.

What to watch is the Pass^8 column rather than the headline average. A model that closes the books correctly six times out of eight has not closed the books. It has produced something a controller still has to check, which is the part the 100 per cent does not cover.