A mid-tier price and a flagship score

Anthropic released Claude Sonnet 5.5 on Monday, six days after its Opus 5.5 flagship, and the company’s own figures put the cheaper model two points behind the expensive one on the evaluation it uses for general knowledge work. On GDPval-AA v2.1, Anthropic reports Sonnet 5.5 at 1,844 against Opus 5.5’s 1,846.

The per-token price has not moved. Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5 and half what Opus 5.5 charges. Cache reads stay at $0.20 per million. Anthropic’s claim is about the total bill rather than the rate: it says the model produces output more than 30 per cent faster and finishes a task for up to 30 per cent less, because it spends fewer tokens and makes fewer tool calls on the way there.

The coding numbers move a long way

On Terminal-Bench 4.0, an agentic coding evaluation, Anthropic reports Sonnet 5.5 at 70.6 per cent against Sonnet 5’s 10.3 per cent — a gap that says as much about how badly the older model handled that particular harness as about the new one. CursorBench 4.0 goes from 34.1 to 55.5 per cent, with Opus 5.5 at 57.8. FrontierCode 1.1 moves from 42.4 to 46.2 per cent, where Opus 5.5 still leads on 54.4. On OSWorld 2.1, the computer-use test, Sonnet 5.5 reports 80.1 per cent to Opus 5.5’s 81.8.

Network switch ports with patch cables plugged into a rack
Inference cost is the argument Anthropic is making for the model. Stock photograph, not one of Anthropic's data centres. Vladimir Srajber · pexels · Pexels License

Every one of those figures is Anthropic’s own, published alongside the model rather than produced by an outside evaluator. None of them has yet been reproduced independently.

The customers Anthropic quotes describe the same shape of improvement in their own terms. Box’s vice-president of AI said the model was 2.4 times faster and used 12 per cent fewer total tokens. Slack reported roughly 14 per cent fewer output tokens. Lovable, whose annualised revenue passed $600m this month, said its coding agent needed about a third fewer tool calls to finish a job. Zendesk said support tickets were processed 20 per cent faster.

Safeguards borrowed from the flagship

Sonnet 5.5 is the first Sonnet model to ship with classifiers meant to stop someone extracting its reasoning to train a rival — a distillation defence Anthropic had previously kept for its largest models. Its cybersecurity safeguards are set at the Opus 5.5 level, and Anthropic says higher-risk requests are routed back to Sonnet 5 rather than answered by the new model. Approved defenders can apply to a Cyber Verification Program for access to the restricted capability.

Two colleagues looking at code on a large monitor in an office
Anthropic says the model finishes tasks with fewer tokens and fewer tool calls. Stock photograph. Mikhail Nilov · pexels · Pexels License

That is a conservative posture to take on a cheap model, and it arrives in a month when agents running on frontier systems have reached live government portals and public code registries without being asked to.

What to watch

The model is on the Claude platform and on AWS, Google Cloud and Microsoft Azure as claude-sonnet-5-5. The figures worth waiting for are the independent ones. Anthropic’s per-task cost claim rests on the model choosing to do less work, and that is exactly the behaviour that varies most between a vendor’s harness and a customer’s codebase. The public coding leaderboards and the intelligence indices will put a number on the difference within days.