Aleph Alpha released Kolibri-1 today: 78.1 billion parameters, 3.46 billion active per token, Apache 2.0 weights, a one-million-token context window, and the pitch of being a sovereign German model. This morning there were two readings of it in circulation. The company’s own technical report says it lies “on the Pareto frontier of quality and serving cost” among the open models it evaluated. The first independent write-up I found, from Trending Topics, is headlined “No Match for the Open-Weight Leaders.”
Both can be true, because they are answering different questions. The question is how much of the headline survives once you read the report’s own tables, which are unusually forthcoming about where the model loses.
The 0.8-point lead
The company’s headline claim is that Kolibri scores best among models with about 3B active parameters. The report puts the margin in plain text: Kolibri scores 0.8 points above Qwen3.5 35B-A3B in English and 1.0 in German, “at 3.46B active parameters against 3B.” On the model card’s post-trained table the overall English scores are 75.5 for Kolibri and 74.7 for Qwen3.5, with Qwen3.8 27B, a dense model, on top at 80.2 (Figure 1). The report says as much: Qwen3.8 27B “has the highest overall aggregates,” and the three models that span the frontier are Qwen3.5, Kolibri and Qwen3.8.
Two caveats apply. First, Kolibri activates about 15 percent more parameters than the 3B model it edges (3.46 / 3, my arithmetic), so the margin is not entirely free. Second, every number here is Aleph Alpha’s own, run on a benchmark suite Aleph Alpha chose, with baselines served at each vendor’s documented settings. That is normal for a launch, and the report is open about method, but it is one party’s measurement. I found no independent score today; Trending Topics notes it is not yet listed on Artificial Analysis and estimates, from the comparison models, an Intelligence Index of roughly 15 to 20. That is an estimate by the writer, not a measurement.
What the 4.4 percent buys
The percentage in the report’s abstract is the real structure of the model: 3,457,573,120 active of 78,103,074,560 total, or 4.43 percent, with 384 routed experts per layer and six selected per token. The routed-expert ratio is a smaller 1.56 percent; the active share is higher because attention, the shared expert, the router and the output head run on every token.
What a ratio like that buys is decode speed. Each generated token computes with about 3.5 billion parameters, not 78 billion. The report’s efficiency chart makes that explicit: its horizontal axis is estimated decode throughput per GPU, measured with vLLM on a single node of eight B200s, FP8 weights and an FP8 KV cache, at high concurrency. It calls that “a proxy for efficiency in workloads where decoding dominates,” and on that axis a sparse model is supposed to win.
What it doesn’t price is residence. The model card gives the footprint plainly: about 78 GB of FP8 weights, with a minimum of two A100 80 GB cards, two H100s, or one H200, B200 or B300. A dense 27B model at FP8 is roughly 27 GB of weights, about a third of that (one byte per parameter, weights only, my arithmetic). Which of those is “cheaper to serve” depends on who’s serving. A cloud tenant keeping eight B200s saturated with requests is optimizing tokens per GPU-second, and Kolibri’s architecture is built for that. An agency or manufacturer running one box for a few hundred employees is mostly buying memory and idling the arithmetic units, and there the 4.4 percent does less than the headline implies. The report is aimed at that second reader, with “public administration, industry, and aerospace” named in its abstract, so the cost axis it chose is a slightly odd fit for its own audience. I’d read the Pareto claim as true on the axis stated and untested on the one a regulated buyer might actually care about.
The million tokens
The same pattern shows up in the context length. The model card’s table lists 1,048,576 tokens, with a recommendation to serve at 262,144 or fewer; the report puts the native length at 262,144, the length of the final training phase, and treats the rest as extrapolation. It does this on purpose. Only the sliding-window layers (512 preceding tokens, four of every five layers) carry rotary position embeddings; the full-attention layers carry none. Because nothing in the model is tied to a position scale, the model card says the context can be extended “in principle to arbitrary lengths,” with no rescaling.
Does it work? The report’s own RULER numbers for the base model are more candid than the headline. Up to 256k tokens, Kolibri trails Qwen3.5 35B-A3B Base, Nemotron 3 Nano 30B-A3B Base and Gemma 4 26B-A4B Base. At 256k it scores 69.8 against 72.1 for Nemotron and 80.1 for Qwen3.5. Past its trained window it degrades more slowly, and at 1M it scores 63.2 against 58.5 and 57.5 (Figure 2). The report says exactly that: it “degrades the least and has the best score at 1M tokens.” Gemma has no score at 256k and beyond, because its window is shorter than the test.
So the one-million-token claim is real in the narrow sense the report supports: of the three models scored at that length, Kolibri holds up best, at a lower starting point. Whether 63 on a synthetic retrieval average is a score you’d build a workflow on is a separate question the benchmark can’t answer.
Sovereign, and by what definition
The report’s framing sentence is that “teams in Germany developed Kolibri end to end and train it on infrastructure in Germany and Finland,” and the abstract says the company controls “the full model-development process, including data, architecture, training infrastructure, post-training, and evaluation.” The weights are Apache 2.0, so a deploying organization can run them on its own hardware. That is a coherent definition of sovereign: who controls the process, and who controls the deployment.
It is not a claim of a closed supply chain, and the report doesn’t pretend it is. Its post-training section names the models used to generate and regenerate training data: GLM-5.2 and GLM-5.3 from Z.ai, and Alibaba’s Qwen3.8-27B, with the stated reason that different models “have different strengths” and “different values and biases, which we consciously select or avoid.” A separate section uses Mistral-Nemo to rewrite German web documents. So the teachers include Chinese-lab models and a French one, and the weights that come out are German-controlled. I find that more interesting than contradictory: sovereignty in 2026 turns out to mean control of the pipeline’s decisions, not the absence of foreign models in it. It is also worth knowing that the company agreed in September to merge with Cohere, the Canadian lab, per SiliconANGLE, with the combined firm to operate under the Cohere brand once regulators clear it. None of that undoes the release, but it does mean “sovereign” describes the company as it stands today.
What I’d want before believing the strong version
The weak version of the claim, that this is a competitive sparse model with an unusually clean long-context story and a transparent report, looks well supported by the document itself. The strong version needs two things the document can’t supply: a third party running the same suite on the post-trained weights, and a cost curve with memory on one axis and throughput on the other. Until then the margin is 0.8 points in English over the nearest 3B-active model, the cost axis leaves out the part a small deployer pays for, and the best-documented fact in the release is that its authors told you where it loses.
References
- Aleph Alpha (2026). Kolibri-1. Hugging Face model card, 3 October 2026. Source of parameter counts, hardware requirements, context-length and benchmark tables.
- Aleph Alpha (2026). Kolibri: A Sovereign European Model on the Pareto Frontier. Technical report, 3 October 2026. Source of the quoted passages on the Pareto frontier, the 0.8-point margin, RULER results, and post-training data generation. The “arbitrary lengths” phrase is from the model card (reference 1).
- Aleph Alpha (2026). Kolibri has landed: a sovereign open-weight model. Company blog, 3 October 2026.
- Steinschaden, J. (2026). Aleph Alpha’s Sovereign A.I. Model Kolibri Is No Match for the Open-Weight Leaders. Trending Topics, 3 October 2026. The Intelligence Index range is the article’s own estimate.
- SiliconANGLE (2026). Cohere and Aleph Alpha agree to merge in reported $20B deal. 16 September 2026. The valuation is an anonymously sourced press report and is not used here.