On the morning of 3 October, Germany's Reunification Day, Heidelberg-based Aleph Alpha opened a quiet download page: Kolibri. A 78.1B-parameter language model with open weights , published on Hugging Face under the Apache 2.0 license. That license lets anyone use the model commercially, run it on their own servers, and keep their data in-house. The presenter's point in the video is exactly this: an AI you own instead of a rented service you depend on. Aleph Alpha announced the release on its own blog, and the timing carries an unmistakable sovereignty message.
What makes Kolibri efficient is its mixture-of-experts design. Inside the model sit 384 small expert networks plus a router that picks the best 6 for every token, alongside 1 shared generalist running on each layer. Of the 78.1B total parameters, only about 3.46B — 4.4% — go active per token. Knowledge accumulated at the 78B scale thus runs at the compute cost of a far smaller model. The stack has 50 layers: 10 run full attention while 40 settle for a 512-token sliding window, which sharply lightens the memory load on long documents.
On context, two numbers come together: a validated ceiling of 1,048,576 tokens and a recommended daily limit of 262,144 tokens. Reaching the ceiling needs special serving flags, and four reasoning-effort levels can be set per request. A purpose-built 128k-token vocabulary reaches 4.90 bytes per token on German text, meaning 11% fewer pieces than generic tokenizers. I checked the serving facts against TheFrontier compilation and the table matches the official card line for line.
Numbers and caveats
Aleph Alpha's published score table looks strong at first glance: 96.9% on AIME 2025 reasoning, 96.0% on AIME 2026, 84.3% on GPQA Diamond, 85.9% on LiveCodeBench and 92.7% on HumanEval+ for coding. The comparison runs against open-weight rivals from Qwen, Nemotron and Mistral, with German sub-splits included. The critical footnote must not be skipped: every figure comes from the vendor's own harness, with no independent replication yet. I took the vendor-scorecard warning from the Traictory analysis, which also shows rivals ahead on retail and function-calling rows.
The real target is not the leaderboard but regulated sectors: public administration, industrials and aerospace. Training ran on infrastructure in Germany and Finland under European law, personal data was redacted before training, and the EU AI Act, the code of practice and data-protection rules are part of the design. A training-content summary was published in the EU template. The idea is plain: institutions that cannot send sensitive documents outside run the model inside their walls and keep data sovereignty . Reading this positioning with the official announcement puts the table in its right place.
The video's three workflows start by turning member-engagement data into a content calendar: thirty days of watch data go into the model, the three hottest topics and the three highest-drop-off titles are found, and a four-week plan comes out. The second workflow is a 24/7 onboarding assistant; the whole tutorial archive, call records and roadmaps are loaded in, and a new member gets a personal first-week plan even at midnight. Both rest on grounded response behavior: when the evidence is not in the documents, the model says it lacks information instead of inventing. The presenter calls this principle critical for member-facing work.
Three workflows and the agent side
The third workflow qualifies leads: form answers and CRM records plug into the model, and every prospect gets a follow-up message tuned to their situation instead of generic copy. Underneath sits native tool calling : the model talks to calendars, forms and databases and acts like a small agent. I verified the agent-call details with TheTechBriefs review, which also notes Kolibri beating rivals clearly on the banking-assistant row while trailing on overall function calling. A strong showcase, then, on a narrow front.
An honest setup warning is owed here: 3.46B active parameters sound small, yet the full file loads into memory — about 78 GB in FP8, double that at full precision. Minimum setup is two 80 GB accelerators or a single H200, B200 or B300, and serving needs the company's vLLM plugin. The card model name garbled in the video's spoken hardware line should defer to the official card's list. For those without the hardware, one developer announced free hosted access at launch, with more community options expected to follow.
How it was trained, why it matters
The training scale is serious: 768 B200 accelerators ran 21 days of pre-training over 20T tokens, followed by 3.44T tokens of mid-training and a 201B-token long-context stage. The blog text says 200B where the card says 201B; the card rules. The team first validated the pipeline on the 30B Kolibri Origin with hundreds of comparative trials, then passed Kolibri through the same line. The German share is 21.3% with translated text at 6%, since organic German data was preferred over translation's cultural fingerprint. Details fill a 189-page technical report.
A corporate layer sits on top of all this: per the CellCog summary, the company signed a combination agreement with Cohere on 16 September, pending regulatory approval. A team releasing serious open weights while heading into consolidation says much about the cost of building an independent AI line in Europe. I took the Cohere combination note from the CellCog summary, which also records over a thousand downloads on day one. The card's last line should not be forgotten either: human review is requested before acting on outputs. Until independent tests land, the table is a promise and the file is the fact.
| Measure | Value |
|---|---|
| Total / active | 78.1B total, 3.46B active |
| Context | 1M validated, 262k advised |
| License and setup | Apache 2.0, 78 GB FP8 |
Key moments
AI commentary
"Ownership matters more than the leaderboard here; the numbers belong to the vendor, the file to everyone. Judgment before independent tests is premature."
AI assessment
The strongest objection sits at the source of the scorecard: the numbers come from Aleph Alpha's own harness and no rival lab has replicated them. Even the company's own compilation shows a dense 27B rival ahead on the English overall average, proving sparsity does not win on every front. Trailing rows on telecom and function calling balance the picture. Waiting a few weeks for community verification beats launch-day enthusiasm.
The missing list is not short: the 1M-token window opens only with special flags, with 262k recommended; the Apache 2.0 grant covers weights and configuration files while training method and code stay out. The blog and the card telling different stories on the German data share casts a small shadow on the transparency claim. The human-review requirement also says the model should not sit alone as a decision-maker.
The presenter is a digital stand-in with a paid community and ready-made workflow playbooks behind the narrative, so the praise deserves filtering. Still, the practical takeaway is clear: regulated teams holding H200-class hardware should download the file and test inside their walls; teams without it should watch hosted options and wait for independent measurements. The ownership idea charms, but its bill arrives in the hardware column.
Sources
6 links; 2 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Julian Goldie SEO
- @aleph-alpha.com Aleph Alpha — Kolibri Has Landed
Also cited by: Germany's Free Open Model: What Kolibri-1 Actually Offers
- @thefrontier.dev TheFrontier — Kolibri-1 open weights
- @traictory.com Traictory — Kolibri vendor scorecard
Also cited by: Germany's Free Open Model: What Kolibri-1 Actually Offers
- @cellcog.ai CellCog — Kolibri benchmarked
- @thetechbriefs.com TheTechBriefs — Kolibri 78.1B MoE
kolibri 78b · aleph alpha · open weights · mixture of experts · data sovereignty · rag · agents