Three values per weight

PrismML released Ternary Bonsai 2 27B on 17 September, a compressed build of Alibaba’s Qwen3.8 27B that stores each weight as one of three values — minus one, zero or plus one — with FP16 scaling applied group by group. The company puts the result at 1.76 effective bits per weight and a total footprint of 5.9GB, more than nine times smaller than the full-precision original.

The compression claim is the whole story, so the retention number carries the weight. Across a 20-benchmark suite covering reasoning, maths, coding, instruction following, vision and agentic tool use, PrismML reports an aggregate of 83.9 for the ternary build against 85.4 for the full-precision Qwen3.8 27B — 98.2 per cent retention. These are the company’s own measurements, published alongside the release, not an independent evaluation.

They are also an improvement on its own previous attempt. PrismML says the first Bonsai 27B retained about 95 per cent; the gap it has closed this time is roughly three points of the base model’s score.

A desktop computer case with the side panel removed
Illustration: the company reports 46.8 tokens per second on an Apple M5 Max. Anete Lusina · pexels · Pexels License

Where 5.9GB puts a model

Quantisation research mostly lives in papers. The practical claim here is about placement: at 5.9GB, a 27-billion-parameter model with a 262,000-token context and text-and-image input fits in consumer memory, alongside an operating system and whatever else is running.

PrismML reports up to 143 tokens per second on an Nvidia GeForce RTX 5090 and 46.8 tokens per second on an Apple M5 Max — laptop-class hardware in the second case. It also gives an energy figure of 0.714 milliwatt-hours per token on an RTX 4090, which it says is about 40 per cent more efficient than full-precision 8-billion-parameter models.

The comparison worth holding onto is that last one. A ternary 27B that costs less energy per token than a dense 8B is not competing with the frontier; it is competing with the small models people currently run locally because nothing bigger fits.

Close-up of the cooling fans on a graphics card
Illustration: PrismML reports up to 143 tokens per second on an RTX 5090. Matheus Bertelli · pexels · Pexels License

The licence is the other half

Bonsai 2 27B is released under Apache 2.0, which permits commercial use without a separate agreement.

That is worth noting because the base model’s lineage has been moving the other way. Alibaba released Qwen-Image-2.1 on Sunday under its Qwen Research License, which restricts use to non-commercial purposes, after shipping the earlier Qwen-Image under Apache 2.0. A derivative that carries a more permissive licence than its family’s recent releases is an unusual position, and it depends entirely on the terms of the specific base checkpoint PrismML built on.

What to watch

The retention figure needs independent confirmation. A 20-benchmark aggregate published by the party selling the compression is a starting point, not a result, and ternary schemes have historically degraded unevenly — holding up on short, well-structured tasks and falling away on long-horizon agentic work, which is exactly the part of the suite hardest to summarise into one number.