Forty per cent the labs cannot see

Artificial Analysis published version 4.2 of its Intelligence Index on 4 September, and the headline change is not a score. Privately held-out test sets now account for 40 per cent of the index weighting, double their share in version 4.1. The held-out portion is made up of AA-Briefcase, AA-Omniscience and the solutions to CritPt, and Artificial Analysis says the share will grow again in version 5.

The reasoning is straightforward. An index built from public benchmarks can be optimised against, whether deliberately or through contamination of training data, and a model that has effectively seen the test is not being measured. Keeping the answers private is the only defence a third-party evaluator has.

Artificial Analysis describes v4.2 as an interim release that pulls parts of the planned version 5 forward to keep the index in step with the frontier.

What went in, and what came out

Two evaluations were added. AA-Briefcase is Artificial Analysis’s own agentic knowledge-work set, built with industry experts around realistic multi-step projects and graded by a mix of rubric and pairwise scoring across task success, analytical quality and presentation. GDP.pdf, built by Surge AI, tests single-turn professional document reasoning across 100 PDFs in ten domains, requiring a model to assemble evidence spread over 4,592 pages of text, tables, charts, footnotes and exclusions.

An analyst reviewing charts on a large monitor
AA-Briefcase grades agentic knowledge work on task success, analytical quality and presentation. Leeloo The First · pexels · Pexels License

One evaluation came out. GPQA Diamond, the graduate-level science reasoning set that has appeared on frontier model cards for two years, was dropped — Artificial Analysis calls it an exceptional evaluation that is now saturated. Alongside the swap, AA-LCR moved to version 1.1 with its own grading prompt and corrected answer keys, and GDPval-AA v2 and AA-Briefcase were given improved sampling and a re-anchored Elo scale so ratings stay stable as new models arrive.

The order after the rebuild

On the rebuilt index, Anthropic’s Claude Fable 5.1 leads. OpenAI’s GPT-6 Astra sits second, with what Artificial Analysis reports as roughly a four-point gain over the previous OpenAI entry. Meta is the third-placed lab, ahead of SpaceXAI, Moonshot’s Kimi, Z.AI and Google.

These are independent measurements, not vendor-reported figures: Artificial Analysis runs each evaluation itself through the models’ public APIs. That is the whole reason a number like this carries weight when a lab’s own benchmark table does not.

Printed charts and a spreadsheet on a desk
Scores from version 4.2 are not comparable with version 4.1, because the components and weightings changed. Max Vakhtbovych · pexels · Pexels License

Why version numbers matter here

One consequence deserves stating plainly. Because the component set and the weightings changed, a v4.2 score cannot be compared with a v4.1 score. A model that appears to have gained or lost ground between the two versions may simply have been measured differently — and the removal of a saturated benchmark mechanically compresses the top of the range.

What to watch is version 5, where Artificial Analysis has said the held-out share rises further. The trade it is making is transparency for integrity: the more of the index the public cannot inspect, the harder it is to game and the harder it is to audit.