Germany’s AI lab picks Reunification Day for its first big open model

Aleph Alpha released Kolibri on 3 October, the Day of German Unity, and put the full weights on Hugging Face under the Apache 2.0 licence. It is the Heidelberg company’s first open-weight release at this scale. A predecessor, Kolibri Origin, finished pre-training on 11 June and was never published.

Kolibri is a mixture-of-experts transformer with 78.1 billion parameters in total and 3.46 billion active for any given token: 384 experts, six of which fire at once. Aleph Alpha said it trained the model on 20 trillion tokens, filtered down from more than 200 trillion tokens of raw data, and that 21.3% of the pre-training mix was German. Translated text makes up 6% of the whole, which the company said it kept low because translated text tends to carry the cultural fingerprint of its source language.

The context window goes up to one million tokens, though the longest sequence the model was actually trained on is 262,144. Its knowledge cutoff is 18 June 2026 for both English and German.

Network cables plugged into a rack-mounted switch
Aleph Alpha is aiming Kolibri at buyers who cannot send internal documents to an outside inference service. Illustration. Brett Sayles · pexels · Pexels License

Jurisdiction, not the frontier

Aleph Alpha is not claiming to have caught the frontier labs. It is claiming jurisdiction. The company said Kolibri was built by its own teams in Germany and trained on infrastructure in Germany and Finland, under European and German law with no foreign control, and that it designed the model against the EU AI Act, the General-Purpose AI Code of Practice and the GDPR from the start.

The pitch is aimed at public administration, industrials and aerospace: buyers who cannot send internal documents to a third-party inference service. Running 3.46 billion parameters per token is small enough to serve on-premise, which is the point of the architecture.

The model is also trained to decline. Aleph Alpha used abstention data and a protocol it calls Merlin-Arthur so that Kolibri says it does not know when the answer is not in the supplied context, a behaviour it says customers ask for and that it tracks as a metric through training.

The numbers, and who produced them

Every published Kolibri score is Aleph Alpha’s own. On its figures the model reaches 96.9 on AIME 2025, 90.0 on a German translation of AIME 2026, 84.3 on GPQA Diamond and 85.9 on LiveCodeBench v6. The company compared it with Alibaba’s Qwen3.6-35B-A3B, Nvidia’s Nemotron 3 Super and Mistral Small 4, and said Kolibri matches models with up to four times its active parameter count. The lead is not uniform: on a banking agent benchmark it reports 38.1 against 10.6 for Qwen3.6-35B-A3B, but on BFCL v4 it reports 61.4 against Qwen’s 67.2.

A hand writing mathematical equations on a whiteboard
Every Kolibri benchmark figure published so far comes from Aleph Alpha itself. Illustration. cottonbro studio · pexels · Pexels License

The comparison set is itself contested. Trending Topics noted that all three rivals Aleph Alpha chose date from the spring, and that Kolibri is not measured against Qwen3.8, GLM-5.3 or Kimi K3. On its reading of the Artificial Analysis Intelligence Index, Kolibri would land behind at least twenty other open-weight models. Aleph Alpha’s answer is that public benchmarks do a poor job of describing what its customers need. No independent evaluation of Kolibri has been published yet.

Three months from 30 billion to 78 billion

The release is also an argument about speed. Kolibri Origin had 30.6 billion total parameters, a 65k context window and 7.51 trillion training tokens. Kolibri finished pre-training on 11 September, three months later. In between, Aleph Alpha tripled the expert count, replaced its routing algorithm, swapped full attention for a sliding window of 512 tokens with full attention every fifth layer, and taught the model four levels of reasoning effort.

Pre-training ran 21 days and hit 38 unplanned interruptions from hardware faults and dropped connections, roughly one per 10,000 GPU-hours, all of which the company said its pipeline handled without a person stepping in.

What to watch is whether an independent leaderboard confirms any of this, and whether German public bodies actually deploy the model. Aleph Alpha’s internal customer-proxy evaluations are the only evidence so far that its specialisation works.