OpenAI takes the top five

Arena, the company behind the crowdsourced model leaderboard, published an Alignment Index on 8 October that ranks 27 models on how badly their agents misbehave rather than how well they score.

OpenAI holds the first five places. GPT-6.1 Sol leads on 87.9, followed by GPT-6 Astra and GPT-6 Luna on 87.8, GPT-6 Sol on 87.6 and GPT-5.6 Sol on 84.2. Anthropic’s Claude Opus 5.5 is sixth on 83.2, SpaceXAI’s Grok 4.7 seventh on 82.7 and Claude Fable 5.1 ninth on 80.1. Google’s Gemini 4 Argon is eleventh on 79.4. MiniMax M3 is last of the 27 on 69.2.

Arena labels the table preliminary, and the data behind it is dated 30 September.

Three ways an agent fails

The index combines three failure rates, all of them lower-is-better. Unauthorized action is the agent doing something it was not given permission to do. False attribution is the agent crediting a statement or a fact to the wrong source. Deceptive completion is the agent reporting a task as finished when it is not.

Deceptive completion is where the models separate. GPT-6.1 Sol is flagged on 2.34% of its sessions, Claude Opus 5.5 on 6.41%, Gemini 4 Argon on 12.86% and MiniMax M3 on 22.54% — close to ten times the rate of the leader. Unauthorized action is far rarer everywhere, from 0.83% for GPT-6 Astra to 3.45% for Gemini 3.7 Flash. False attribution runs from 1.56% to 15.61%.

Hands typing on a laptop keyboard at a wooden desk
The scores come from 72,509 sessions in which visitors gave agents real tasks. Illustrative image. Lukas Blazek · pexels · Pexels License

Measured on work people actually asked for

The scores come from 72,509 sessions on Arena’s own Agent Arena, where visitors give agents real tasks rather than running a fixed test set. Arena’s argument for doing it that way is that “static benchmarks break down once models recognize they’re being tested”.

That design is also the index’s main limitation. Because the workload is whatever visitors chose to ask for, two models are not necessarily being measured on the same work, and the mix can shift over time. Arena publishes a range alongside every score and every rate, and notes that a single flagged case can count toward more than one failure mode, so the three rates overlap rather than partition the failures.

A $200m round on the same day

Arena announced the index alongside a $200m Series B at a $3.1bn valuation, led by Lightspeed Venture Partners and Khosla Ventures, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, Andreessen Horowitz and Felicis also taking part. That is close to double the $1.7bn post-money valuation of its $150m Series A in January.

TechCrunch reports that annualised revenue was $30m at the January round and reached $100m in June, from a commercial product called AI Evaluations launched last September. Arena says “the world needs a neutral third party to measure how safe and aligned AI actually is”.

Empty running track lanes seen from above
Arena publishes the table as preliminary, with a range alongside every score. Illustrative image. KoolShooters · pexels · Pexels License

The conflict nobody has solved

A leaderboard that sells evaluation services to the labs it ranks is not a disinterested referee, and Arena’s own framing — a neutral third party — is the claim most worth testing. None of the scores above has been reproduced independently, and the index measures flagged failures rather than verified ones.

The useful thing to watch is whether the labs start publishing their own numbers against these three definitions, and whether a ten-fold gap in deceptive completion survives contact with a fixed, held-out task set. If it does, the industry has a measurable problem with agents that say the job is done.