What changed

Artificial Analysis published version 4.3 of its Intelligence Index on 7 September, two days after shipping version 4.2. The firm describes both as interim releases pulled forward from a planned v5, and the pace is the point: the index is being revised faster than the models it ranks are shipping.

Two evaluations moved. Terminal-Bench went from v2.1 to v4.0, which the maintainer describes as harder terminal tasks spanning several domains. And a benchmark called τ³-Banking was dropped in favour of AutomationBench-AA, a business workflow automation test with a private test set, inheriting the same 5% weighting.

The category weights did not move: Agents 30%, General 30%, Coding 20%, Scientific Reasoning 20%. What did move is how much of the index a lab cannot see. Evaluations with private test sets now carry 45% of the total, up from 40% in v4.2 — which was itself double the share in v4.1.

Handwritten notes covering an office whiteboard
Terminal-Bench moved from v2.1 to v4.0 in the revision. Stock image. K · pexels · Pexels License

The ranking

On v4.3, Claude Fable 5.1 and GPT-6 Astra are tied in first place on 53. Claude Opus 5 follows on 51, Claude Fable 5 on 50, Meta’s Muse Spark 1.3 on 48 and GPT-5.6 Sol on 47.

That is a change from v4.2, published two days earlier, where Fable 5.1 led alone. Artificial Analysis attributes the shift to where the two models differ: Fable 5.1 scores higher on AA-Briefcase and SciCode, while Astra scores higher on Terminal-Bench v4.0 and on the newly added AutomationBench-AA. Swap in a harder terminal test and an agentic automation set, and the gap closes.

The index now runs on ten evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.

The number underneath the tie

The more useful figure is not the tie but the cost of it. Artificial Analysis puts GPT-6 Astra at $3.26 per task against Claude Fable 5.1’s $7.63 — 57% less to reach the same measured score.

That is the kind of comparison a buyer can act on, and it is also the kind that needs its caveats stated. Cost per task depends on how many tokens a model spends thinking, which varies by task mix; a different set of ten evaluations would produce a different ratio.

Colleagues meeting around a table with open laptops
Two revisions in two days changed the leaderboard without any model changing. Stock image. Christina Morillo · pexels · Pexels License

Reading an index that moves this often

These are independent measurements, not vendor-reported scores, which is what makes them worth reading at all. But two revisions in two days is a reminder of what a composite index actually is: a weighted opinion about which capabilities count. Change the constituent tests, and the ranking changes without any model changing.

The rising share of private test sets is a direct response to that fragility. If nearly half the index cannot be trained against, a score is harder to manufacture. It also means the public cannot audit that half — the trade-off Artificial Analysis has chosen, and one worth naming.

What to watch

Whether v5 arrives with the same constituents or another reshuffle, and whether the Astra–Fable tie survives it. A first place that changed hands on a benchmark swap rather than a model release is a fact about the index as much as about the models.