What was released

The Allen Institute for AI published Olmo-core 3 on Thursday, an open training framework for large mixture-of-experts language models, together with a technical report and an interactive demo.

A mixture-of-experts model routes each token to a small subset of its parameters, so total capacity can grow without the cost of every forward pass growing with it. That is why nearly every frontier model is now sparse. It is also why training one is harder than training a dense model: the routing has to be load-balanced across devices, and the communication pattern it creates is the thing that decides whether the hardware is busy or idle.

The numbers

In a preliminary test on eight Nvidia B300 GPUs, Ai2 says a 47-billion-parameter mixture-of-experts processed 52,000 tokens per second per GPU under the new stack, against 19,400 under its earlier implementation built on PyTorch’s FSDP — about 2.7 times the throughput.

Cooling fans inside a computer case
The institute reports 2.7 times the throughput of its previous implementation. Illustrative photo. Andrey Matveev · pexels · Pexels License

The capacity figures are the other half. Ai2 grew the expert pool from 8 to 128 while keeping four experts active per token, which took total parameters from 4.6 billion to 47 billion at roughly 3.2 billion active per token, and cost less than 5 per cent of training throughput.

At scale, the institute reports benchmarking a 1.2-trillion-parameter configuration with 58.36 billion parameters active per token across 512 B300 GPUs, with a peak observed throughput of 858 TFLOP/s per GPU, and a further test at 2.38 trillion parameters using DeepEP v2. Enabling the MXFP8 number format where it helped most gave about 21 per cent more throughput than the BF16 baseline, and brought peak active memory down from 103 GiB to 95 GiB.

These are the institute’s own measurements of its own framework, published with the release.

Why the infrastructure matters more than the weights

Open-weight models are now common. Open training infrastructure capable of a trillion-parameter sparse run is not, and it is the part that determines whether a university group or a small lab can do anything other than fine-tune someone else’s checkpoint.

A technician installing hardware in a server rack
The largest configuration tested spanned 512 GPUs. Illustrative photo. panumas nikhomkhai · pexels · Pexels License

Ai2 has been consistent about this: its Olmo line has shipped data, code and intermediate checkpoints rather than weights alone. Olmo-core 3 extends that to the layer that is usually the most jealously kept, because distributed training performance is where a lot of a frontier lab’s practical advantage actually sits.

The caveat is hardware. These numbers come from B300s, and a stack tuned for the newest Nvidia parts is less useful to a group running older accelerators — which is most groups outside the well-funded ones.

What to watch is whether anyone outside Ai2 reproduces the throughput figures on their own cluster. An open framework is only as good as its second user.