Reflection opens Beam, and leads with the running cost
Reflection AI introduced Beam on 5 October, the first open-weight model from a lab founded in 2024 by two former Google DeepMind researchers, Misha Laskin and Ioannis Antonoglou. The weights are not out yet. Reflection said it will publish them under the Apache 2.0 licence later in October, together with a technical report, a model card and tooling for running, evaluating and fine-tuning the model, once red-teaming and evaluation finish. Access before then is by waitlist.
Beam is a text-only sparse mixture-of-experts transformer: 501 billion parameters in total, 23 billion active for any single token. Reflection said it pre-trained the model on 23.8 trillion tokens of web and licensed data in under four weeks on 6,144 Nvidia GB300 NVL72 chips, then ran a reinforcement learning stage of more than 100 million rollouts on 10,500 of the same chips over a further four weeks. The context window was 256,000 tokens during reinforcement learning and was stretched to a million in midtraining.

The claim is parity with GLM 5.2 at a fraction of the compute
Reflection’s pitch is not that Beam is the strongest open model available. It is that Beam reaches scores “comparable to GLM-5.2 while using 3–4× less inference compute” — a cost argument rather than a capability one. Zhipu AI’s GLM 5.2 carries roughly 250 billion more parameters than Beam does.
Every benchmark figure published so far is Reflection’s own, and none has been reproduced independently. On the company’s table, Beam scores 80.1 on Terminal-Bench v2.1 against 81.0 for GLM 5.2, 86.6 for Alibaba’s Qwen 3.8-Max and 88.3 for Moonshot AI’s Kimi K3. On GPQA Diamond it scores 90.5, against 91.2 for GLM 5.2, 92.6 for Qwen 3.8-Max and 93.5 for Kimi K3. On AIME 2026 it reaches 97.8 against GLM 5.2’s 99.2. On the hard split of SWE-Bench Pro v2 it scores 77.2, well behind Kimi K3 at 84.3 and Qwen 3.8-Max at 88.2.
Read straight, the table shows Beam about a point behind GLM 5.2 across the board and clearly behind the two larger Chinese models. That is close to what Reflection says it shows: the company calls Beam “competitive with” GLM 5.2 and says it is “approaching” Qwen 3.8-Max on coding and agentic work.

What it cost to get here
Reflection has raised about $4.7bn, including from Nvidia, Sequoia Capital and Lightspeed Venture Partners, at a $25bn pre-money valuation, TechCrunch reported. It has signed more than $7bn in compute agreements with SpaceX and Nebius for access to Nvidia GB300 chips through 2029.
The company’s framing is geopolitical: Chinese labs hold the open-weight frontier, and a Western lab should hold part of it. On the evidence Reflection has published, Beam does not take that frontier back. What it offers is a model that can be downloaded, modified and self-hosted under a permissive licence, at an inference bill its maker says is a third to a quarter of the nearest comparable open model’s.
The weights and the technical report are the thing to watch for later this month. Independent runs of Terminal-Bench and SWE-Bench Pro will settle whether the efficiency claim survives outside Reflection’s own harness.