Skip to content

Kolibri-1 ships with open weights and an Apache licence: sovereignty stops at the GPU

Published on 5 October 2026

Sala con dos armarios de servidores iluminados tenuemente y, en primer plano, unas llaves sobre un plano doblado encima de una mesa.

Aleph Alpha has released Kolibri-1 with open weights under Apache 2.0. Before you read that as "at last, a European model I can run at home": no, you won't be running this at home. Here are the numbers and where the catch is.

What happened

On 3 October 2026, Aleph Alpha pushed Kolibri-1 to Hugging Face. It's a mixture-of-experts model — only part of the network fires for each token — with these figures:

  • 78,103,074,560 total parameters, of which 3,457,573,120 are active per token. Big-model capacity, small-model compute cost.
  • 50 layers, 384 experts per layer (1 shared, 6 routed), attention at a 4:1 ratio between sliding-window and GQA layers.
  • FP8 weights in 128×128 blocks, FP8 KV cache too. Memory footprint: around 78 GB for the weights alone.
  • Native context of 262,144 tokens, validated up to 1,048,576. They themselves recommend staying at or below 262,144 for serving efficiency.
  • German and English only. Knowledge cutoff of 18 June 2026 for both.
  • Apache 2.0, with no usage restrictions bolted on top.

The training details are unusually specific: 20 trillion pre-training tokens split roughly 62.5% English, 23.9% German and 13.6% code; another 3.44 trillion in mid-training and 201 billion in the long-context phase. All of it on 768 B200 GPUs over 21 days, 392,000 GPU-hours for pre-training, 6.4e23 FLOPS. Estimated energy use: about 950 MWh, data-centre overhead included.

The scores they publish on the card: 84.3 on GPQA Diamond, 80 on MMLU-Pro, 96 on AIME 2026, 66.4 on SWE-bench Verified, 21.5 on HLE and 27.7 on Terminal-Bench 2.1.

Why it matters

Cost per token and cost of entry are two different things, and here they diverge badly. With 3.46 billion active parameters, each token is cheap to compute. But memory isn't negotiable: all 78 GB have to be resident even though only a slice works at a time. The stated minimum is two 80 GB A100s, two H100 SXM5, one H200, one B200 or one B300. Recommended: two H100 SXM5, two H200s, one B200 or one B300. In practice: this doesn't fit on a 4090 or your laptop, and if you plan to host it yourself the budget starts with the metal.

The licence does change the conversation. Apache 2.0 means you deploy it where you like, you control the version, and nobody can pull it or rewrite the terms mid-contract. That is real operational sovereignty, and it's different from "open" licences with usage clauses you have to read through a PDF before signing with a client.

It is not drop-in on vLLM. You need the aleph-alpha-inference package, which ships its own plugin, and the server has to be started with specific parsers for reasoning mode and tool calls. That pins your serving layer to a particular vLLM version. If you run several models behind one server, budget for keeping that dependency current.

Long context is the serious argument. 262,144 tokens natively, extendable, is a lot for internal documentation work, structured extraction and search over your own material. That's where a model like this pays for the hardware — not playing general chatbot.

What doesn't change

Every evaluation figure comes from the model card and is flagged as unverified by third parties. They're self-reported; treat them as a hypothesis you still have to confirm against your own data.

German and English only. If your use case is in Spanish, this isn't your model — and to be fair, nobody claimed otherwise: the tokeniser is tuned to German word structure, and covering two languages instead of twenty is an explicit decision.

That 27.7 on Terminal-Bench says plainly that as a loose agent in a terminal it isn't something you delegate to unwatched. The card owns it: they place the model in workflows where a person reviews output before anything acts on it, and on the advisory side of decision support, never as the decider.

And the energy figure excludes supervised fine-tuning, reinforcement learning and ablation models. It's an honest estimate, not the full total.

Our take

The part of this release I like most isn't a benchmark: it's the model card. They state the GPUs, the days, the megawatt-hours, the minimum hardware, the biases they expect and what you shouldn't use it for. That's worth more than half a point on MMLU-Pro. I'm tired of launches that amount to a bar chart and a post on X, and then you discover the VRAM requirement three hours into the fight. Aleph Alpha has also signed the EU code of practice for general-purpose AI, which is exactly the sort of boring paperwork that saves you meetings with a client's legal department.

That said, the word "sovereignty" grates a little. The model was trained on 768 B200s. European sovereignty ends where Nvidia's catalogue begins, and no licence fixes that. What you do get, and it isn't small, is deployment sovereignty: the weights are yours, you choose where they run, and nobody deprecates your endpoint with three months' notice. On a three-year contract that weighs more than five benchmark points.

I think the bilingual bet will age poorly. It makes technical sense — a tokeniser tuned for German genuinely saves tokens — but the market compares against models that handle thirty languages acceptably, and a single-language edge evaporates within a year. What doesn't evaporate is holding the weights.

If this came to me for evaluation, I wouldn't slot it in as a generalist model. I'd put it behind one concrete, measurable job — long German documentation, structured extraction, retrieval over in-house material — and I'd measure two things before quality: real GPU occupancy and latency at the context length you'll actually use, not the brochure maximum. And I'd compare at equal hardware, not equal parameter count, which is exactly where mixture-of-experts models fool the eye.

This is precisely what happens when someone tells us "this can't leave our network": ten minutes in, the conversation has stopped being about the model and is about GPUs, who updates the inference server, and who gets out of bed if the service falls over on a Sunday. The model is almost never the hard part.

Any questions, tell me and we'll look at it.

Best, Vicente.

Source: The Decoder · Via: The Decoder

Did reading this raise a question?

Ask us. We answer even if you never become a client.