On October 3, 2026, Aleph Alpha released a new large language model called Kolibri. The MoE (Mixture-of-Experts) Transformer, which works in English and German, has a total of 78.1 billion parameters, though only 3.46 billion of them are active for each token. The model's full weights were published on the Hugging Face platform under the Apache 2.0 license — meaning it can be downloaded, used commercially, fine-tuned, and deployed on-premise.
The company says it timed the release for the Day of German Reunification (October 3):
«On the Day of German Reunification, we are releasing our new model: Kolibri.» — Aleph Alpha's official blog
Kolibri is the product of the automated model-training pipeline the company calls its "Model Factory." The same pipeline was previously tested on the Kolibri Origin prototype: a smaller model with 30.6 billion parameters and a 65,000-token context window that was never released publicly. According to the company, in the three months between the two releases the parameter count grew from 30 to 78 billion, context length from 65,000 to 1 million tokens, and training volume from 7.5 to 20 trillion tokens. To obtain those 20 trillion tokens, the pipeline processed more than 200 trillion raw tokens — filtering them, removing duplicates, and curating them.
Technical Specifications
Kolibri targets the balance between serving cost and quality thanks to its MoE architecture: in the company's own comparison, the model sits on the quality-cost "Pareto frontier" for English and German — meaning there is no better model at the same price, and no cheaper model at the same quality. The key technical facts:
- 78.1 billion parameters in total, 3.46 billion active parameters per token;
- context length of up to 1 million tokens (the longest trained length — 262,144 tokens);
- bilingual, 128K tokenizer; 21.3% of pre-training tokens are in German;
- knowledge cutoff — June 18, 2026;
- release date on the Hugging Face model card — October 3, 2026, license — Apache 2.0.
The architecture details are also open: all 50 layers are MoE (with one shared expert), 384 experts in total, 6 of them active per token. The attention mechanism combines a sliding window of size 512 with full attention in every fifth layer. Pre-training started at a base length of 16,384 tokens.
The company places special emphasis on bilinguality: the tokenizer was purpose-built for German and English, with little translated text used (only 6%), since translations often carry the cultural imprint of the source language. As the blog puts it, the result is not a model "trained in English that read a bit of German," but a model bilingual by design.
Benchmark Results: The Company's Own Claims
All benchmark scores cited in the Kolibri announcement are Aleph Alpha's own measurements; they have not yet been confirmed by independent sources. In the company's table, Kolibri scored: AIME 2025 — 96.9, AIME 2026 — 96.0, GPQA Diamond — 84.3 (81.3 on the German variant). For comparison, the same table lists Qwen3.6-35B-A3B (AIME 2025 — 84.6), Nemotron 3 Super 120B-A12B (91.7), and Mistral Small 4 119B-A6B (79.8).
According to the blog's conclusion, Kolibri matches models with four times as many active parameters, including Nemotron 3 Super, in mathematics, coding, grounding, and long-context tasks. However, the company admits that public benchmarks do not fully reflect clients' domain-specific needs — which is why it developed its own internal evaluation sets for sectors like the German public sector, aviation, manufacturing, and automotive. In these internal tests, the model's scores grew steadily across post-training stages, as shown in charts, though this data also rests on the company's own reporting and has not been independently verified.
On the scale of training, the blog writes: pre-training ran for 21 days on 20 trillion tokens across 768 NVIDIA B200 GPUs; Kolibri Origin finished training on June 11, 2026, and Kolibri on September 11, 2026. According to the company, 38 unplanned interruptions occurred during those 21 days — due to hardware failures or connection drops — and the pipeline resolved all of them automatically without human intervention: the cluster restarted training on other nodes, with the run resuming from a checkpoint at most 250 steps back.
A "Sovereign" Model: Compliance With European Law
Aleph Alpha is positioning Kolibri as a "sovereign" model: it was developed in Germany, trained on infrastructure in Germany and Finland, and built under European and German law. The company says the model was constructed with the EU AI Act, the code of practice for general-purpose AI, and GDPR requirements in mind, with particular attention to copyright issues.
Target sectors are regulated industries: public administration, industry, and aerospace. In the company's statement, sovereignty combines two dimensions: how the model is built and how it is handed to clients. The model's small, efficient size lets clients run it on-premise without sending internal data to third-party inference services. The blog also describes the "Merlin-Arthur" protocol: the model is trained to say "I don't know" when the context does not support an answer — a feature the company says clients requested. The company stresses that the entire supply chain — from data collection and curation through pre-training, post-training, and final evaluation — is transparent.





