Aleph Alpha Kolibri is a new open-weight language model from the Heidelberg AI company Aleph Alpha, released on 3 October 2026, the Day of German Unity. Kolibri, German for hummingbird, is an English and German mixture-of-experts (MoE) model with 78.1 billion parameters, of which only 3.46 billion are active for any one token. The full weights are on Hugging Face under the Apache 2.0 licence, and Aleph Alpha says the model handles up to 1,048,576 tokens of context. That is the “1M context window” in every headline.
The numbers are real, but the million needs a footnote. Aleph Alpha Kolibri was trained on sequences of up to 262,144 tokens. The jump to a million comes from a design choice that lets it run past its training length, and Aleph Alpha’s own model card recommends staying at or below 262,144 tokens “for serving efficiency and complex tasks”. In Aleph Alpha’s own long-context testing, the base model scores 86.9 at 4,000 tokens and 63.2 at a million.
Two more things frame the launch. The rivals Aleph Alpha benchmarks against mostly date from the spring, and the company is about to be folded into Canada’s Cohere, which signed a definitive merger agreement with it on 16 September.
Below we set out what Aleph Alpha Kolibri is, how its 384-expert design works, what the 1M figure does and does not mean, how it scores against other open-weight models, why it is trained to say “I don’t know”, what it takes to run on your own hardware, what the licence covers, and what the Cohere merger means for it.
Table of contents
- What Aleph Alpha Kolibri Is
- How the Aleph Alpha Kolibri Mixture of Experts Works
- The 1M Context Window: What Aleph Alpha Kolibri Was Trained On
- Aleph Alpha Kolibri Benchmarks Against Other Open-Weight Models
- Grounding: Why Aleph Alpha Kolibri Is Trained to Say “I Don’t Know”
- German First: The Aleph Alpha Kolibri Tokenizer and Training Data
- Running Aleph Alpha Kolibri on Your Own Hardware
- The Licence: What Apache 2.0 Covers for Aleph Alpha Kolibri
- Aleph Alpha Kolibri and the Cohere Merger
- Who Should Look at Aleph Alpha Kolibri
- Aleph Alpha Kolibri FAQ
- References and Further Reading
What Aleph Alpha Kolibri Is
Aleph Alpha Kolibri is a reasoning model with tool calling, built for German and English only. Aleph Alpha describes it as “a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace”. The pitch is that a government department or manufacturer can run it on its own servers, without sending internal documents to a US or Chinese cloud.
The headline numbers
The model card lists 78,103,074,560 total parameters and 3,457,573,120 active per token, which is 4.4% of the model doing work at any moment. It has 50 layers, 384 routed experts per layer plus one shared expert, and six routed experts chosen for each token. Its knowledge cutoff is 18 June 2026 for both languages, and it supports four reasoning levels: none, low, medium and high.
Aleph Alpha posted the launch on X at 08:54 UTC on 3 October: “Small bird, fast wings, Kolibri is here.” The Hugging Face repository had been created the day before, and the inference plugin it needs was tagged v1.0.0 at 07:11 UTC on launch day. By the evening of 4 October, the community had already published 27 derived versions of Aleph Alpha Kolibri on Hugging Face, including 7 GGUF builds and 11 for Apple’s MLX.
Kolibri Origin, the model nobody saw
Aleph Alpha Kolibri is the second model of its family. The first, Kolibri Origin, was a 30.6-billion-parameter model with a 65,536-token window that Aleph Alpha used to test its training pipeline and never released. Origin finished pre-training on 11 June 2026. Kolibri finished on 11 September, three months later.
| Property | Kolibri Origin | Kolibri |
|---|---|---|
| Release | No public release | 3 October 2026 |
| Total parameters | 30.6B | 78.1B |
| Active parameters per token | 3.27B | 3.46B |
| Experts (total / active) | 128 / 8 | 384 / 6 |
| Pre-training tokens | 7.51T | 20T |
| Longest trained length | 65,536 tokens | 262,144 tokens |
| Attention | Full attention, all layers | 512-token sliding window, full attention every fifth layer |
| Tokenizer vocabulary | 96,000 | 128,000 |
| Reasoning modes | One | None, low, medium, high |
Built in Germany and Finland
Aleph Alpha says its teams built the model in Germany and trained it on infrastructure in Germany and Finland, “under European and German law, with no foreign control”. Pre-training ran on 768 Nvidia B200 GPUs for 21 days. Across those three weeks the job hit 38 unplanned interruptions, roughly one per 10,000 GPU-hours, and the pipeline restarted itself each time without anyone stepping in.
How the Aleph Alpha Kolibri Mixture of Experts Works
A mixture-of-experts model splits its feed-forward layers into many small sub-networks and uses a router to send each token to a few of them. The rest sit idle for that token. You get the knowledge capacity of a large model with the compute cost of a small one. The catch, which the Aleph Alpha Kolibri model card states plainly, is memory: “the full model must be held in memory even though only part of it is active at any time.”
384 experts, six at a time
Each of the 50 layers in Aleph Alpha Kolibri holds 384 routed experts with a hidden size of 512, plus one shared expert that every token passes through. A sigmoid router picks the top six routed experts per token. Aleph Alpha tested fewer, wider experts and found that many small ones performed better. It also tripled the expert count from Kolibri Origin’s 128 while cutting the number chosen per token from eight to six.
Keeping 384 experts evenly loaded is hard, because routers like to send most tokens to a favoured few. Aleph Alpha’s answer is what it calls exact quantile balancing. Moonshot’s Kimi K3 introduced quantile balancing with an estimated quantile; Aleph Alpha says it computes the quantile exactly at a fixed cost and that the exactness improves both load balance and quality.
Sliding windows and ten full-attention layers
Only 10 of the 50 layers in Aleph Alpha Kolibri look at the whole context. The other 40 use a sliding window of 512 preceding tokens plus the current one. This four-to-one pattern keeps memory and compute for those 40 layers fixed no matter how long the prompt gets, which is what makes very long contexts affordable to serve. Attention uses 48 query heads and 4 key-value heads.
Why 78B and not 123B
Aleph Alpha says bigger kept getting better in its tests, all the way from 32B to 123B. It stopped at 78B because of serving cost. On two Nvidia H100s with 256,000-token requests, the 123B version could serve 3 users at once, while the 78B Aleph Alpha Kolibri served 18 and decoded 28% faster.
Concurrent 256K-token requests on two H100 GPUs (Aleph Alpha’s figures)
The bars are scaled against 18, so the 123B bar is 3 ÷ 18 = 17% as long. Six times the concurrency for a model that scored somewhat lower is the trade Aleph Alpha chose, and it tells you who the model is for: organisations paying for their own GPUs.
The 1M Context Window: What Aleph Alpha Kolibri Was Trained On
The context window is where the headline and the documentation part ways. Aleph Alpha Kolibri was trained in three stages, each on longer sequences than the last, and none of them reached a million tokens.
| Stage | Sequence length | Tokens | Time on 768 B200s |
|---|---|---|---|
| Pre-training | 16,384 | 20T | 21 days |
| Mid-training | 65,536 | 3.44T | 5 days |
| Long-context extension | 262,144 | 201B | 13 hours |
| Served by extrapolation | 1,048,576 | None | Not trained |
Trained to 262,144 tokens
The model card says it plainly: Aleph Alpha Kolibri “was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length.” The config file that ships with the weights sets max_position_embeddings to 262,144. Out of the box, that is the limit.
How it reaches 1,048,576
The trick is in where the position information lives. Aleph Alpha Kolibri applies rotary position embeddings only inside the 40 sliding-window layers, where no token ever looks more than 512 places back. The 10 full-attention layers have no positional encoding at all. Because nothing in the model has learned “position 300,000” as a special case, Aleph Alpha says the context “can be extended beyond that length without any position scaling, in principle to arbitrary lengths”. To serve a million tokens you pass --max-model-len 1048576 and override the position limit when you start the server.
What the long-context tests show
Aleph Alpha tested the base model on RULER, a standard long-context benchmark, at lengths up to a million. Its own 189-page technical report sums it up as useful capability “up to 1M tokens, with task-dependent degradation”.
Kolibri Base RULER score by context length (model card, 0 to 100)
Each bar’s width is the score itself on a 0 to 100 scale. Two things stand out. The decline from 256K to 1M is gentle, 6.6 points, so the extrapolation really does hold up. And at a million tokens Aleph Alpha Kolibri beats the two rivals tested there, Nemotron 3 Nano at 58.5 and Qwen3.5 35B-A3B at 57.5. The bigger drop, though, comes earlier: from 86.9 at 4K to 67.9 at 128K.
The tech report also says the losses are uneven. At a million tokens Kolibri Base beats Nemotron 3 Nano on needle-in-a-haystack, variable tracking and question answering, but not on word extraction, and its needle score falls as the context grows. So a million-token prompt will work, but anything that depends on finding one fact in a huge document should be tested before you rely on it.
What a million tokens holds
Aleph Alpha measures its tokenizer at 4.58 bytes of English web text per token. At that rate, 1,048,576 tokens is about 4.8 million bytes, or 4.8 MB of plain English text. At roughly six bytes per English word including the space, that is in the region of 800,000 words. In practice, Aleph Alpha’s advice to stay at or under 262,144 tokens still leaves room for about 200,000 words, which covers most contract bundles and policy archives.
Aleph Alpha Kolibri Benchmarks Against Other Open-Weight Models
Aleph Alpha ran every comparison with its own open-source eval framework, with Kolibri at its highest reasoning setting. These are vendor numbers, not independent ones. Artificial Analysis had not listed Aleph Alpha Kolibri when Trending Topics checked on 3 October.
| Benchmark | Kolibri 3B active | Qwen3.6 35B-A3B | Qwen3.5 35B-A3B | Nemotron 3 Super 120B-A12B | Mistral Small 4 119B-A6B | GPT-OSS 120B |
|---|---|---|---|---|---|---|
| Overall (EN) | 75.5 | 71.4 | 74.7 | 73.0 | 63.1 | 72.3 |
| Overall (DE) | 70.8 | 67.3 | 69.8 | 67.9 | 61.4 | 70.2 |
| GPQA Diamond (EN) | 84.3 | 83.4 | 83.8 | 78.0 | 74.7 | 76.4 |
| AIME 2026 (EN) | 96.0 | 91.0 | 92.1 | 90.4 | 83.1 | 90.8 |
| LiveCodeBench v6 | 85.9 | 82.5 | 77.8 | 82.0 | 71.2 | 87.5 |
| SWE-Bench Verified | 66.4 | 73.8 | 71.6 | 60.2 | 60.8 | No score |
| TerminalBench 2.1 | 27.7 | No score | 39.7 | 39.7 | 21.0 | 29.2 |
| BFCL v4 tool calling | 61.4 | 67.2 | 70.5 | 61.0 | 58.0 | 57.3 |
| Tau2-Bench Telecom | 94.7 | 99.1 | 97.7 | 68.1 | 41.5 | 73.1 |
| AA-LCR long context | 68.3 | 69.7 | 66.3 | 67.0 | 52.3 | No score |
“No score” means the model card reports none, usually because a prompt exceeded that model’s window or it could not call tools in that harness.
Where it leads
On maths and science reasoning, Aleph Alpha Kolibri leads the mixture-of-experts field it chose. It scores 96.9 on AIME 2025, 96.0 on AIME 2026 and 84.3 on GPQA Diamond. Its English overall average of 75.5 edges Qwen3.5 35B-A3B at 74.7 and Nemotron 3 Super at 73.0, a model with roughly three and a half times as many active parameters. The biggest gap is on Tau3-Bench Banking, an agent task, where Kolibri scores 38.1 against 15.5 for Nemotron 3 Super and 10.6 for Qwen3.6.
Where it trails
Coding agents and tool use are weaker. On SWE-Bench Verified, Kolibri’s 66.4 trails both Qwen models. On TerminalBench 2.1 it scores 27.7 against 39.7 for Qwen3.5 and Nemotron 3 Super. On the multi-turn part of the BFCL tool-calling test it scores 47.5 against 62.7 for GLM-4.7 Flash and 65.2 for GLM-4.5 Air. If your main use is an autonomous coding agent, Aleph Alpha Kolibri is not the leader here.
English overall score, post-trained models (model card, 0 to 100)
Bar widths equal each score on a 0 to 100 scale. The “8x” for Qwen3.8 27B is its 27 billion active parameters divided by Kolibri’s 3.46 billion, which comes to 7.8.
The comparison set is from the spring
The Austrian site Trending Topics made the sharpest point about the launch: the three rivals in Aleph Alpha’s headline charts, Qwen3.6-35B-A3B, Nemotron 3 Super and Mistral Small 4, all date from the spring. Newer open-weight models such as GLM-5.3, Kimi K3 and Xiaomi’s MiMo-V2.6-Pro are missing. Trending Topics estimated that Kolibri would score around 15 to 20 on the Artificial Analysis Intelligence Index, behind at least 20 other open-weight models.
That needs one correction. Trending Topics wrote that Kolibri “is not compared with any” of the Qwen3.8 family. The full table in the model card does include Qwen3.8 27B, but greys it out as a dense model that activates several times as many parameters per token. On that row Qwen3.8 27B leads on most tests, 80.2 to 75.5 in English and 79.9 to 70.8 in German. Aleph Alpha Kolibri’s claim is about quality per unit of serving cost, not raw quality.
Small differences between blog and card
The launch blog and the model card disagree on a few figures. The blog gives FRAMES as 71.2 for Kolibri, the card 73.0. The blog says 21.3% of pre-training tokens are German, while the card’s data table adds up to 23.9% German. The tech report says “more than 20%”, which both meet. None of this changes the picture, but it is a reason to quote the model card when the numbers matter.
Grounding: Why Aleph Alpha Kolibri Is Trained to Say "I Don't Know"
Most AI training rewards a guess, because a guess is sometimes right and an abstention never scores. Aleph Alpha argues that for a regulated customer “a model that knows to abstain is the difference between a pilot and a deployment”. So Aleph Alpha Kolibri was trained to decline when the documents it is given do not contain the answer.
Merlin, Morgana and Arthur
The method is Aleph Alpha’s Merlin-Arthur protocol. Arthur is the model being trained. Merlin edits a document so the right answer becomes easier to find. Morgana strips out the evidence and tries to tempt Arthur into a confident guess. Arthur cannot tell which one he faces, so the only winning strategy is to read the context carefully and abstain when the evidence is gone. The game generates training examples automatically from documents Aleph Alpha already holds.
The numbers
On the public set of the AA-Omniscience test, Aleph Alpha Kolibri avoids a wrong answer on 44.0% of items, against 15.0% for Kolibri Origin. On the RGB benchmark it holds back when the documents lack an answer 85.6% of the time. Aleph Alpha’s own “M/A grounding score”, which it says certifies a lower bound on how much of an answer came from the document, puts Kolibri at 0.23, where Origin and several rivals score 0.
AA-Omniscience non-hallucination rate, public set (model card, % of items not answered wrongly)
Bar widths equal the percentage. Aleph Alpha’s launch blog highlights the jump from Origin and the lead over Nemotron 3 Super. The same model card table shows Qwen3.6 35B-A3B, a model Aleph Alpha uses as a headline rival, hallucinating less on this test than Aleph Alpha Kolibri does.
The trade-off
Declining to answer has a cost. On AA-Omniscience accuracy, Kolibri gets 14.8% of items right, lower than every other mixture-of-experts model in its table except Origin. On RGB’s fact-check task, which asks the model to spot and fix errors, it scores 34.0 against 74.0 for Qwen3.5 and 90.0 for Nemotron 3 Super. A model tuned to stay inside the documents is a good fit for document search and a weaker one for open-ended questions from memory.
German First: The Aleph Alpha Kolibri Tokenizer and Training Data
Most open-weight models treat German as one language among dozens. Aleph Alpha Kolibri supports two, by design. The model card calls it “a deliberate choice of depth over breadth”, and the blog says the result is “bilingual by design, not an English model that has read some German”.
UniBPE and compound words
German builds long compound nouns, and standard tokenizers chop them at odd points. Aleph Alpha trained a new tokenizer method it calls UniBPE, a mix of the two common approaches, BPE and Unigram. Its example is “Bundessozialgerichtes” (of the Federal Social Court). Aleph Alpha Kolibri splits it as Bundes, sozial, gericht, es. GPT-5, Qwen3.8 and Gemini all split it as Bund, ess, oz, ial, gericht, es, which cuts “Bundes” in the wrong place.
| Tokenizer | Vocabulary | German web, bytes per token | English web, bytes per token |
|---|---|---|---|
| Kolibri | 128,000 | 4.90 | 4.58 |
| GPT-5 | 200,019 | 4.35 | 4.67 |
| Qwen3.5 to Qwen3.8 | 248,077 | 4.17 | 4.47 |
| Gemini | 262,144 | 4.13 | 4.49 |
| Tekken (Mistral, Nemotron, Apertus) | 131,072 | 4.03 | 4.45 |
| GLM 5.3 | 154,856 | 3.93 | 4.61 |
| DeepSeek V4 | 129,280 | 3.72 | 4.59 |
| Kimi K3 | 163,586 | 3.28 | 4.62 |
More bytes per token means fewer tokens for the same text. Aleph Alpha’s tech report puts the German saving at 11.2% fewer tokens than GPT-5, the best of the nine other tokenizers it tested. The arithmetic checks out: 4.35 ÷ 4.90 = 0.888, so the same German text needs about 11% fewer tokens.
Finding four trillion German tokens
Small test models told Aleph Alpha that about 20% German was the best mix, which meant finding roughly 4 trillion German tokens for a 20-trillion-token run. After filtering, open German datasets supplied only 390 billion. Aleph Alpha built its own German Common Crawl pipeline, retuned so it no longer threw out long administrative words, which yielded 1.3 trillion tokens. Rewriting German documents in new styles added about a trillion more. Kolibri saw the resulting 2.4-trillion-token pool about 1.8 times on average.
Why it matters for cost
Fewer tokens per document means lower bills, shorter waits and more room in the context window. A German public body sending long legal texts through Aleph Alpha Kolibri pays for about 11% fewer tokens than with a GPT-5-class tokenizer, before any difference in model price. The model also reasons in German rather than only answering in it, which makes its reasoning easier for German-speaking staff to check.
Running Aleph Alpha Kolibri on Your Own Hardware
The default download is an FP8 checkpoint of about 78 GB. A separate BF16 repository holds full-precision weights at about 156 GB. Here is what Aleph Alpha says each needs.
| Version | Weights in memory | Minimum | Recommended |
|---|---|---|---|
| Kolibri-1 (FP8) | About 78 GB | 2x A100 80 GB, 2x H100 SXM5, 1x H200, 1x B200 or 1x B300 | 2x H100 SXM5, 2x H200, 1x B200 or 1x B300 |
| Kolibri-1-BF16 | About 156 GB | 4x A100 80 GB, 4x H100 SXM5, 2x H200, 1x B200 or 1x B300 | 4x H100 SXM5, 2x H200, 2x B200 or 1x B300 |
Those are data-centre GPUs. The 3.46 billion active parameters make each token cheap to compute, but all 78 billion still have to sit in GPU memory. For smaller machines, community builds quantised to 2, 3, 4, 6 and 8 bits appeared on Hugging Face within two days, though Aleph Alpha has not tested them. Our local LLM hardware guide covers what consumer cards can realistically hold.
The vLLM plugin
Aleph Alpha Kolibri does not run on stock vLLM. You need Aleph Alpha’s aleph-alpha-inference package, which adds the model code plus Kolibri-specific parsers for reasoning and tool calls. It is Apache 2.0 too, and also ships as a container image. Its package file pins vLLM to the 0.29 series, with the note that “each release supports exactly one vLLM minor version”. Expect to wait for a plugin update before moving to a newer vLLM.
Settings that matter
The model card recommends a temperature of 1.0, top_p of 0.97 and top_k of 128. Reasoning effort is set per request through the chat template: none, low, medium or high. Higher effort means more thinking tokens, so it costs more and answers more slowly. The server exposes an OpenAI-compatible API, so existing client code works by changing the base URL and model name.
Where it fits among open-weight options
Aleph Alpha Kolibri sits at the small, cheap-to-serve end of the field. A much larger mixture-of-experts model such as DeepSeek V4.1 Flash at 552B total needs far more memory. Our guide to open-weight AI models in 2026 compares the wider field.
The Licence: What Apache 2.0 Covers for Aleph Alpha Kolibri
Apache 2.0 is one of the most permissive licences there is. You can use the weights commercially, modify them and redistribute them, with no user caps and no field-of-use limits.
Weights yes, recipe no
The model card adds a limit worth reading. The licence “only applies to the weights and configuration files published in this repository” and “does not extend to underlying code, model architecture, parameter settings or any training method”. Aleph Alpha keeps all rights to its training code and methods. So Aleph Alpha Kolibri is open-weight, not open-source in the fuller sense that projects such as OLMo use, where data and training code are also released.
Guidance, not conditions
The responsible-use section asks people to avoid prohibited practices under Article 5 of the EU AI Act and activities “related to military or nuclear applications”. Because the weights are under Apache 2.0, these are requests, not licence terms. One launch report, from Crypto Briefing, described the model as pitched at “public administration, industry, aerospace, and defense”. Aleph Alpha’s own blog names public administration, industrials and aerospace, and does not mention defence.
Where it sits under the EU AI Act
Aleph Alpha says it built Aleph Alpha Kolibri with the EU AI Act, the GPAI Code of Practice and GDPR in mind, and it is a Code of Practice signatory. Training used 6.4 × 10^23 floating-point operations, according to the model card. The Act presumes systemic risk above 10^25, so Kolibri used about 6.4% of that threshold. Under Article 53(2), open-source models below that line are excused two documentation duties, but still need a copyright policy and a public summary of training content. Aleph Alpha has published that summary on the European Commission’s template.
The energy bill
The model card estimates training energy at 950 MWh, including data-centre overhead. That covers pre-training, mid-training and long-context extension, which used 392,000, 90,000 and 10,000 GPU-hours, so 492,000 in total. It excludes fine-tuning and reinforcement learning. Dividing 950,000 kWh by 492,000 GPU-hours gives about 1.9 kWh per B200 GPU-hour, all overheads included.
Aleph Alpha Kolibri and the Cohere Merger
Aleph Alpha Kolibri lands at an odd moment for the company. Aleph Alpha is being absorbed by Cohere, the Toronto-based enterprise AI company, and the release may be one of the last under the Aleph Alpha name.
| When | What happened |
|---|---|
| April 2026 | Cohere and Aleph Alpha announce a planned combination, with Schwarz Group backing |
| 11 June 2026 | Kolibri Origin finishes pre-training |
| 11 September 2026 | Kolibri finishes pre-training |
| 16 September 2026 | Definitive business combination agreement signed |
| 3 October 2026 | Kolibri released under Apache 2.0 |
| Later in 2026 | Deal expected to close, subject to regulatory approval |
The deal
We covered the plan when it was announced in our piece on the Cohere Aleph Alpha merger. The definitive agreement, reported by The Next Web on 16 September, keeps the Cohere name, with headquarters in Berlin and Toronto and more than 1,000 staff. Heidelberg stays open as a research centre. Schwarz Group, owner of Lidl and Kaufland, plans to invest €500 million, and the combined company is reportedly valued at about $20 billion.
What it means for Aleph Alpha Kolibri
The weights are out under Apache 2.0, and that cannot be taken back. Anyone who downloads Aleph Alpha Kolibri today can keep using it whatever happens to the brand. What is less certain is the roadmap. Aleph Alpha co-founder Samuel Weinbach is due to become Cohere’s chief research officer and co-chief executive Ilhan Scheer its chief operating officer. Cohere sells its own Command models, and neither company has said whether a Kolibri 2 would ship under that name.
Who Should Look at Aleph Alpha Kolibri
Strong fit
Public bodies and regulated firms in German-speaking countries that need a model on their own servers are the obvious audience. So are teams building document search and question answering where a wrong answer costs more than no answer. And anyone who needs cheap long-context serving will find the sliding-window design useful, provided they stay near the 262,144-token native window.
Weaker fit
Teams that need the best open model for coding agents should look at Qwen3.6 or Qwen3.5 first, going by Aleph Alpha’s own table. Anyone working in French, Spanish or any language other than German and English is outside what Aleph Alpha Kolibri was built for. And if your workload depends on recalling facts from training rather than from supplied documents, its low accuracy on knowledge tests matters.
Test it before you trust the 1M figure
If the million-token window is the reason you are interested, run your own documents through it at the length you need. Aleph Alpha’s own numbers show the model works at that length, with a lower score, and its own advice is to stay at 262,144 tokens or below for demanding work. For most document workloads that is still a very large window.
Aleph Alpha Kolibri FAQ
What is Aleph Alpha Kolibri?
Aleph Alpha Kolibri is an open-weight English and German language model released on 3 October 2026. It is a mixture-of-experts model with 78.1 billion total parameters and 3.46 billion active per token, with reasoning and tool calling, under the Apache 2.0 licence.
Does Aleph Alpha Kolibri really have a 1M context window?
It can be served at 1,048,576 tokens, and Aleph Alpha has tested it there. It was trained only up to 262,144 tokens, which is its native length, and the model card recommends staying at or below that for demanding tasks. Its RULER score falls from 69.8 at 256K to 63.2 at 1M.
What hardware does Aleph Alpha Kolibri need?
The FP8 version needs about 78 GB of GPU memory: at minimum two A100 80 GB or two H100 cards, or one H200, B200 or B300. The BF16 version needs about 156 GB. Community quantisations are smaller but untested by Aleph Alpha.
Is Aleph Alpha Kolibri open source?
The weights are under Apache 2.0, so you can use, change and redistribute them commercially. The training code, data and methods are not released, so it is open-weight rather than fully open-source.
How does Aleph Alpha Kolibri compare with Qwen and Mistral?
In Aleph Alpha’s tests it beats Qwen3.6 35B-A3B, Qwen3.5 35B-A3B and Mistral Small 4 on overall English and German scores and on maths. It trails both Qwen models on coding agents and tool use. Newer models such as GLM-5.3, Kimi K3 and the larger Qwen3.8 releases were not in its main comparison.
What happens to Kolibri after the Cohere merger?
The released weights stay available under Apache 2.0 regardless. The combined company will operate as Cohere once regulators approve the deal, with Heidelberg kept as a research centre. Neither company has said what the next Kolibri release will be called.
References and Further Reading
Kolibri Has Landed: A Sovereign Open-Weight Model (Aleph Alpha)
Aleph-Alpha/Kolibri-1 model card (Hugging Face)
Aleph-Alpha/Kolibri-1-BF16 (Hugging Face)
Kolibri: A Sovereign European Model on the Pareto Frontier (technical report, PDF)
Kolibri public summary of training content (PDF)
aleph-alpha-inference vLLM plugin (GitHub)
Merlin-Arthur protocol paper (arXiv)
RULER: What’s the Real Context Size of Your Long-Context Language Models? (arXiv)
Aleph Alpha’s Sovereign A.I. Model Kolibri Is No Match for the Open-Weight Leaders (Trending Topics)
Aleph Alpha releases open-weight Kolibri with 1M context (TestingCatalog)
Aleph Alpha Releases Kolibri (MarkTechPost)
Cohere and Aleph Alpha sign, with headquarters in Berlin and Toronto (The Next Web)
Cohere and Aleph Alpha agree to merge in reported $20B deal (SiliconANGLE)
EU AI Act Article 53: Obligations for providers of general-purpose AI models
EU AI Act Article 51: Classification of general-purpose AI models with systemic risk
The General-Purpose AI Code of Practice (European Commission)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.