What changed
Artificial Analysis published version 4.3 of its Intelligence Index on 7 September, two days after shipping version 4.2. The firm describes both as interim releases pulled forward from a planned v5, and the pace is the point: the index is being revised faster than the models it ranks are shipping.
Two evaluations moved. Terminal-Bench went from v2.1 to v4.0, which the maintainer describes as harder terminal tasks spanning several domains. And a benchmark called τ³-Banking was dropped in favour of AutomationBench-AA, a business workflow automation test with a private test set, inheriting the same 5% weighting.
The category weights did not move: Agents 30%, General 30%, Coding 20%, Scientific Reasoning 20%. What did move is how much of the index a lab cannot see. Evaluations with private test sets now carry 45% of the total, up from 40% in v4.2 — which was itself double the share in v4.1.

The ranking
On v4.3, Claude Fable 5.1 and GPT-6 Astra are tied in first place on 53. Claude Opus 5 follows on 51, Claude Fable 5 on 50, Meta’s Muse Spark 1.3 on 48 and GPT-5.6 Sol on 47.
That is a change from v4.2, published two days earlier, where Fable 5.1 led alone. Artificial Analysis attributes the shift to where the two models differ: Fable 5.1 scores higher on AA-Briefcase and SciCode, while Astra scores higher on Terminal-Bench v4.0 and on the newly added AutomationBench-AA. Swap in a harder terminal test and an agentic automation set, and the gap closes.
The index now runs on ten evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.
The number underneath the tie
The more useful figure is not the tie but the cost of it. Artificial Analysis puts GPT-6 Astra at $3.26 per task against Claude Fable 5.1’s $7.63 — 57% less to reach the same measured score.
That is the kind of comparison a buyer can act on, and it is also the kind that needs its caveats stated. Cost per task depends on how many tokens a model spends thinking, which varies by task mix; a different set of ten evaluations would produce a different ratio.

Reading an index that moves this often
These are independent measurements, not vendor-reported scores, which is what makes them worth reading at all. But two revisions in two days is a reminder of what a composite index actually is: a weighted opinion about which capabilities count. Change the constituent tests, and the ranking changes without any model changing.
The rising share of private test sets is a direct response to that fragility. If nearly half the index cannot be trained against, a score is harder to manufacture. It also means the public cannot audit that half — the trade-off Artificial Analysis has chosen, and one worth naming.
What to watch
Whether v5 arrives with the same constituents or another reshuffle, and whether the Astra–Fable tie survives it. A first place that changed hands on a benchmark swap rather than a model release is a fact about the index as much as about the models.