A tenth of the price

Anthropic released Claude Haiku 5.5 on 7 October, calling it the cheapest, fastest and most capable small model it has shipped.

The pricing is the news. For prompts up to 100,000 tokens, input costs $0.10 per million tokens and output $0.50. Haiku 4.5 charged $1.00 and $5.00 — ten times as much. Above 100,000 prompt tokens the new model charges $0.50 and $2.50, still half the old rate. Cache reads fall from $0.10 to $0.01, and cache writes from $1.25 to $0.125.

Anthropic puts the average saving at about 75%, with a footnote that resolves the arithmetic: 90% lower for requests under 100,000 tokens, 50% lower above it. Sonnet 5.5’s cache-read price was also halved at launch, from $0.20 to $0.10.

What Anthropic says it scores

Every figure below is vendor-reported, from Anthropic’s own evaluations, and none of it has been independently reproduced.

On OSWorld 2.1, the company’s table gives Haiku 5.5 72.4% against 15.7% for Haiku 4.5 and 48.9% for OpenAI’s GPT-6 Luna. On Terminal-Bench 4.0 it gives 39.2%, against 0.0% for Haiku 4.5 and 16.4% for GPT-6 Luna. On Humanity’s Last Exam without tools it reports 45.9%, up from 10.2%.

A hand using a calculator on a printed sheet of chart types
Input falls from $1.00 to $0.10 per million tokens on prompts under 100,000 tokens. Illustrative image. RDNE Stock project · pexels · Pexels License

Sonnet 5.5 still sits above it on every line — 83.9% on OSWorld, 70.6% on Terminal-Bench — which is the point of the lineup rather than an embarrassment to it. Two customers are quoted: Box reported a score 11 points higher than Haiku 4.5 at roughly half the latency, and AlphaSense reported 0.84 against 0.76 on its own Ask in Document evaluation.

An effort dial on a small model

Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, the control that lets a caller trade spend against how hard the model works on a request. Until now that has been a feature of the larger Claude models.

That matters more than it sounds for the work Haiku is sold for — classification, extraction, routing, summarising the same document shape a million times. A cheap model with a dial can be tuned per queue rather than per application.

Fibre optic cables plugged into a patch panel in a data centre
The model runs on AWS, Google Cloud, Azure and Anthropic's own platform. Illustrative image. Brett Sayles · pexels · Pexels License

The footnote on the price

One detail softens the headline number. Haiku 5.5 uses an updated tokenizer, and Anthropic says it consumes slightly more tokens per task than its predecessor did. A price cut measured per token is not quite a price cut measured per job, and anyone modelling the saving should measure it on their own traffic.

Anthropic also says cybersecurity safeguards on the model are tighter than Haiku 4.5’s but looser than Sonnet 5.5’s, and reports alignment improvements over the previous version without putting a number on them. It does not state the context window.

Where it runs, and what to watch

The model is available now on Amazon Web Services, Google Cloud and Microsoft Azure as well as Anthropic’s own platform.

The thing worth watching is whether the small-model floor keeps dropping. Anthropic has priced Haiku 5.5 into the same territory as OpenAI’s GPT-6 Luna, and the Decisions API OpenAI shipped the same week charges $0.10 per million input tokens with output free. The cheap tier is where the volume is, and both labs now appear to be pricing it as a position rather than a margin.