Fermion Research released Phonon-2 on 30 September, an open-weight English speech recognition model that downloads at 164 megabytes and, the company says, transcribes an hour of audio in roughly twenty seconds on a MacBook Air.
The accuracy claim is the one that makes the size interesting. Fermion reports a 5.21% average word error rate across the seven public English test sets used by the Open ASR Leaderboard, and says every open model that scores better is at least 5.8 times its size. The company has opened a pull request to add the model to that leaderboard, so the figure will shortly be checkable by someone other than its author.
How it got small
Phonon-2 is distilled from Nvidia’s Parakeet TDT 0.6B v3, a 2.5-gigabyte model, and holds each encoder weight at one of five learned levels using about 2.1 bits. That is a 93% reduction in file size against the teacher.

Quantising to roughly two bits per weight is aggressive, and the usual outcome is a model that holds up on clean speech and collapses on accents, noise and overlapping speakers. Whether this one does is exactly what the leaderboard entry will show.
Why it runs anywhere
One file covers Apple silicon, Linux, Windows and Nvidia GPUs. The implementations differ underneath — Apple builds use MLX for GPU acceleration, and the CPU paths use AVX-512, AVX2 or NEON depending on the processor — but the distributed artefact is the same.
It is available through a command line tool installable from PyPI, through Docker, and inside Detta, Fermion’s own macOS dictation app. The weights are published on Hugging Face under CC-BY-4.0, a licence that permits commercial use with attribution.

Why a small model matters here
Speech recognition is one of the few AI workloads where running locally is often the requirement rather than the preference. Medical dictation, legal interviews, newsroom transcription and anything covered by a data residency rule all have the same constraint: the audio should not leave the machine.
A model that fits in 164 megabytes and runs faster than real time on a consumer laptop takes that from an engineering project to a download.
What to watch
Whether the Open ASR Leaderboard entry lands and reproduces the 5.21% figure, and how the model holds up on accented and noisy audio, which the seven public test sets cover only partly.