Aleph Alpha Kolibri is a new open-weight language model from the Heidelberg AI company Aleph Alpha, released on 3 October 2026, the Day of German Unity. Kolibri, German for hummingbird, is an English and German mixture-of-experts (MoE) model with 78.1 billion parameters, of which only 3.46 billion are active for any one token. The full weights are on Hugging Face under the Apache 2.0 licence, and Aleph Alpha says the model handles up to 1,048,576 tokens of context. That is the “1M context window” in every headline.

The numbers are real, but the million needs a footnote. Aleph Alpha Kolibri was trained on sequences of up to 262,144 tokens. The jump to a million comes from a design choice that lets it run past its training length, and Aleph Alpha’s own model card recommends staying at or below 262,144 tokens “for serving efficiency and complex tasks”. In Aleph Alpha’s own long-context testing, the base model scores 86.9 at 4,000 tokens and 63.2 at a million.

Two more things frame the launch. The rivals Aleph Alpha benchmarks against mostly date from the spring, and the company is about to be folded into Canada’s Cohere, which signed a definitive merger agreement with it on 16 September.

Below we set out what Aleph Alpha Kolibri is, how its 384-expert design works, what the 1M figure does and does not mean, how it scores against other open-weight models, why it is trained to say “I don’t know”, what it takes to run on your own hardware, what the licence covers, and what the Cohere merger means for it.

What Aleph Alpha Kolibri Is

aleph alpha kolibri 78b moe 1m context b beer stein with a hinged lid and thumb lever

Aleph Alpha Kolibri is a reasoning model with tool calling, built for German and English only. Aleph Alpha describes it as “a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace”. The pitch is that a government department or manufacturer can run it on its own servers, without sending internal documents to a US or Chinese cloud.

The headline numbers

The model card lists 78,103,074,560 total parameters and 3,457,573,120 active per token, which is 4.4% of the model doing work at any moment. It has 50 layers, 384 routed experts per layer plus one shared expert, and six routed experts chosen for each token. Its knowledge cutoff is 18 June 2026 for both languages, and it supports four reasoning levels: none, low, medium and high.

Aleph Alpha posted the launch on X at 08:54 UTC on 3 October: “Small bird, fast wings, Kolibri is here.” The Hugging Face repository had been created the day before, and the inference plugin it needs was tagged v1.0.0 at 07:11 UTC on launch day. By the evening of 4 October, the community had already published 27 derived versions of Aleph Alpha Kolibri on Hugging Face, including 7 GGUF builds and 11 for Apple’s MLX.

Kolibri Origin, the model nobody saw

Aleph Alpha Kolibri is the second model of its family. The first, Kolibri Origin, was a 30.6-billion-parameter model with a 65,536-token window that Aleph Alpha used to test its training pipeline and never released. Origin finished pre-training on 11 June 2026. Kolibri finished on 11 September, three months later.

PropertyKolibri OriginKolibri
ReleaseNo public release3 October 2026
Total parameters30.6B78.1B
Active parameters per token3.27B3.46B
Experts (total / active)128 / 8384 / 6
Pre-training tokens7.51T20T
Longest trained length65,536 tokens262,144 tokens
AttentionFull attention, all layers512-token sliding window, full attention every fifth layer
Tokenizer vocabulary96,000128,000
Reasoning modesOneNone, low, medium, high

Built in Germany and Finland

Aleph Alpha says its teams built the model in Germany and trained it on infrastructure in Germany and Finland, “under European and German law, with no foreign control”. Pre-training ran on 768 Nvidia B200 GPUs for 21 days. Across those three weeks the job hit 38 unplanned interruptions, roughly one per 10,000 GPU-hours, and the pipeline restarted itself each time without anyone stepping in.

How the Aleph Alpha Kolibri Mixture of Experts Works

aleph alpha kolibri 78b moe 1m context c chest of small drawers with three pulled open

A mixture-of-experts model splits its feed-forward layers into many small sub-networks and uses a router to send each token to a few of them. The rest sit idle for that token. You get the knowledge capacity of a large model with the compute cost of a small one. The catch, which the Aleph Alpha Kolibri model card states plainly, is memory: “the full model must be held in memory even though only part of it is active at any time.”

384 experts, six at a time

Each of the 50 layers in Aleph Alpha Kolibri holds 384 routed experts with a hidden size of 512, plus one shared expert that every token passes through. A sigmoid router picks the top six routed experts per token. Aleph Alpha tested fewer, wider experts and found that many small ones performed better. It also tripled the expert count from Kolibri Origin’s 128 while cutting the number chosen per token from eight to six.

Keeping 384 experts evenly loaded is hard, because routers like to send most tokens to a favoured few. Aleph Alpha’s answer is what it calls exact quantile balancing. Moonshot’s Kimi K3 introduced quantile balancing with an estimated quantile; Aleph Alpha says it computes the quantile exactly at a fixed cost and that the exactness improves both load balance and quality.

Sliding windows and ten full-attention layers

Only 10 of the 50 layers in Aleph Alpha Kolibri look at the whole context. The other 40 use a sliding window of 512 preceding tokens plus the current one. This four-to-one pattern keeps memory and compute for those 40 layers fixed no matter how long the prompt gets, which is what makes very long contexts affordable to serve. Attention uses 48 query heads and 4 key-value heads.

Why 78B and not 123B

Aleph Alpha says bigger kept getting better in its tests, all the way from 32B to 123B. It stopped at 78B because of serving cost. On two Nvidia H100s with 256,000-token requests, the 123B version could serve 3 users at once, while the 78B Aleph Alpha Kolibri served 18 and decoded 28% faster.

Concurrent 256K-token requests on two H100 GPUs (Aleph Alpha’s figures)

Kolibri at 78B total parameters: 18
Tested 123B variant: 3

The bars are scaled against 18, so the 123B bar is 3 ÷ 18 = 17% as long. Six times the concurrency for a model that scored somewhat lower is the trade Aleph Alpha chose, and it tells you who the model is for: organisations paying for their own GPUs.

The 1M Context Window: What Aleph Alpha Kolibri Was Trained On

aleph alpha kolibri 78b moe 1m context d two masted schooner under sail

The context window is where the headline and the documentation part ways. Aleph Alpha Kolibri was trained in three stages, each on longer sequences than the last, and none of them reached a million tokens.

StageSequence lengthTokensTime on 768 B200s
Pre-training16,38420T21 days
Mid-training65,5363.44T5 days
Long-context extension262,144201B13 hours
Served by extrapolation1,048,576NoneNot trained

Trained to 262,144 tokens

The model card says it plainly: Aleph Alpha Kolibri “was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length.” The config file that ships with the weights sets max_position_embeddings to 262,144. Out of the box, that is the limit.

How it reaches 1,048,576

The trick is in where the position information lives. Aleph Alpha Kolibri applies rotary position embeddings only inside the 40 sliding-window layers, where no token ever looks more than 512 places back. The 10 full-attention layers have no positional encoding at all. Because nothing in the model has learned “position 300,000” as a special case, Aleph Alpha says the context “can be extended beyond that length without any position scaling, in principle to arbitrary lengths”. To serve a million tokens you pass --max-model-len 1048576 and override the position limit when you start the server.

What the long-context tests show

Aleph Alpha tested the base model on RULER, a standard long-context benchmark, at lengths up to a million. Its own 189-page technical report sums it up as useful capability “up to 1M tokens, with task-dependent degradation”.

Kolibri Base RULER score by context length (model card, 0 to 100)

4K tokens: 86.9
32K tokens: 76.3
128K tokens: 67.9
256K tokens (trained limit): 69.8
512K tokens: 65.5
1M tokens: 63.2

Each bar’s width is the score itself on a 0 to 100 scale. Two things stand out. The decline from 256K to 1M is gentle, 6.6 points, so the extrapolation really does hold up. And at a million tokens Aleph Alpha Kolibri beats the two rivals tested there, Nemotron 3 Nano at 58.5 and Qwen3.5 35B-A3B at 57.5. The bigger drop, though, comes earlier: from 86.9 at 4K to 67.9 at 128K.

The tech report also says the losses are uneven. At a million tokens Kolibri Base beats Nemotron 3 Nano on needle-in-a-haystack, variable tracking and question answering, but not on word extraction, and its needle score falls as the context grows. So a million-token prompt will work, but anything that depends on finding one fact in a huge document should be tested before you rely on it.

What a million tokens holds

Aleph Alpha measures its tokenizer at 4.58 bytes of English web text per token. At that rate, 1,048,576 tokens is about 4.8 million bytes, or 4.8 MB of plain English text. At roughly six bytes per English word including the space, that is in the region of 800,000 words. In practice, Aleph Alpha’s advice to stay at or under 262,144 tokens still leaves room for about 200,000 words, which covers most contract bundles and policy archives.

Aleph Alpha Kolibri Benchmarks Against Other Open-Weight Models

aleph alpha kolibri 78b moe 1m context e quiz buzzer with a big dome button

Aleph Alpha ran every comparison with its own open-source eval framework, with Kolibri at its highest reasoning setting. These are vendor numbers, not independent ones. Artificial Analysis had not listed Aleph Alpha Kolibri when Trending Topics checked on 3 October.

BenchmarkKolibri 3B activeQwen3.6 35B-A3BQwen3.5 35B-A3BNemotron 3 Super 120B-A12BMistral Small 4 119B-A6BGPT-OSS 120B
Overall (EN)75.571.474.773.063.172.3
Overall (DE)70.867.369.867.961.470.2
GPQA Diamond (EN)84.383.483.878.074.776.4
AIME 2026 (EN)96.091.092.190.483.190.8
LiveCodeBench v685.982.577.882.071.287.5
SWE-Bench Verified66.473.871.660.260.8No score
TerminalBench 2.127.7No score39.739.721.029.2
BFCL v4 tool calling61.467.270.561.058.057.3
Tau2-Bench Telecom94.799.197.768.141.573.1
AA-LCR long context68.369.766.367.052.3No score

“No score” means the model card reports none, usually because a prompt exceeded that model’s window or it could not call tools in that harness.

Where it leads

On maths and science reasoning, Aleph Alpha Kolibri leads the mixture-of-experts field it chose. It scores 96.9 on AIME 2025, 96.0 on AIME 2026 and 84.3 on GPQA Diamond. Its English overall average of 75.5 edges Qwen3.5 35B-A3B at 74.7 and Nemotron 3 Super at 73.0, a model with roughly three and a half times as many active parameters. The biggest gap is on Tau3-Bench Banking, an agent task, where Kolibri scores 38.1 against 15.5 for Nemotron 3 Super and 10.6 for Qwen3.6.

Where it trails

Coding agents and tool use are weaker. On SWE-Bench Verified, Kolibri’s 66.4 trails both Qwen models. On TerminalBench 2.1 it scores 27.7 against 39.7 for Qwen3.5 and Nemotron 3 Super. On the multi-turn part of the BFCL tool-calling test it scores 47.5 against 62.7 for GLM-4.7 Flash and 65.2 for GLM-4.5 Air. If your main use is an autonomous coding agent, Aleph Alpha Kolibri is not the leader here.

English overall score, post-trained models (model card, 0 to 100)

Qwen3.8 27B, dense, about 8x the active parameters: 80.2
Kolibri, 3.46B active: 75.5
Qwen3.5 35B-A3B: 74.7
Nemotron 3 Super 120B-A12B: 73.0
GPT-OSS 120B: 72.3
Qwen3.6 35B-A3B: 71.4
Mistral Small 4 119B-A6B: 63.1

Bar widths equal each score on a 0 to 100 scale. The “8x” for Qwen3.8 27B is its 27 billion active parameters divided by Kolibri’s 3.46 billion, which comes to 7.8.

The comparison set is from the spring

The Austrian site Trending Topics made the sharpest point about the launch: the three rivals in Aleph Alpha’s headline charts, Qwen3.6-35B-A3B, Nemotron 3 Super and Mistral Small 4, all date from the spring. Newer open-weight models such as GLM-5.3, Kimi K3 and Xiaomi’s MiMo-V2.6-Pro are missing. Trending Topics estimated that Kolibri would score around 15 to 20 on the Artificial Analysis Intelligence Index, behind at least 20 other open-weight models.

That needs one correction. Trending Topics wrote that Kolibri “is not compared with any” of the Qwen3.8 family. The full table in the model card does include Qwen3.8 27B, but greys it out as a dense model that activates several times as many parameters per token. On that row Qwen3.8 27B leads on most tests, 80.2 to 75.5 in English and 79.9 to 70.8 in German. Aleph Alpha Kolibri’s claim is about quality per unit of serving cost, not raw quality.

Small differences between blog and card

The launch blog and the model card disagree on a few figures. The blog gives FRAMES as 71.2 for Kolibri, the card 73.0. The blog says 21.3% of pre-training tokens are German, while the card’s data table adds up to 23.9% German. The tech report says “more than 20%”, which both meet. None of this changes the picture, but it is a reason to quote the model card when the numbers matter.

Grounding: Why Aleph Alpha Kolibri Is Trained to Say "I Don't Know"

aleph alpha kolibri 78b moe 1m context f hand crank dynamo with a big flywheel

Most AI training rewards a guess, because a guess is sometimes right and an abstention never scores. Aleph Alpha argues that for a regulated customer “a model that knows to abstain is the difference between a pilot and a deployment”. So Aleph Alpha Kolibri was trained to decline when the documents it is given do not contain the answer.

Merlin, Morgana and Arthur

The method is Aleph Alpha’s Merlin-Arthur protocol. Arthur is the model being trained. Merlin edits a document so the right answer becomes easier to find. Morgana strips out the evidence and tries to tempt Arthur into a confident guess. Arthur cannot tell which one he faces, so the only winning strategy is to read the context carefully and abstain when the evidence is gone. The game generates training examples automatically from documents Aleph Alpha already holds.

The numbers

On the public set of the AA-Omniscience test, Aleph Alpha Kolibri avoids a wrong answer on 44.0% of items, against 15.0% for Kolibri Origin. On the RGB benchmark it holds back when the documents lack an answer 85.6% of the time. Aleph Alpha’s own “M/A grounding score”, which it says certifies a lower bound on how much of an answer came from the document, puts Kolibri at 0.23, where Origin and several rivals score 0.

AA-Omniscience non-hallucination rate, public set (model card, % of items not answered wrongly)

Qwen3.8 27B (dense): 67.3
Qwen3.6 35B-A3B: 56.7
Kolibri: 44.0
Mistral Small 4 119B-A6B: 34.7
GPT-OSS 120B: 23.7
Kolibri Origin: 15.0
Nemotron 3 Super 120B-A12B: 13.9
Qwen3.5 35B-A3B: 11.1

Bar widths equal the percentage. Aleph Alpha’s launch blog highlights the jump from Origin and the lead over Nemotron 3 Super. The same model card table shows Qwen3.6 35B-A3B, a model Aleph Alpha uses as a headline rival, hallucinating less on this test than Aleph Alpha Kolibri does.

The trade-off

Declining to answer has a cost. On AA-Omniscience accuracy, Kolibri gets 14.8% of items right, lower than every other mixture-of-experts model in its table except Origin. On RGB’s fact-check task, which asks the model to spot and fix errors, it scores 34.0 against 74.0 for Qwen3.5 and 90.0 for Nemotron 3 Super. A model tuned to stay inside the documents is a good fit for document search and a weaker one for open-ended questions from memory.

German First: The Aleph Alpha Kolibri Tokenizer and Training Data

Most open-weight models treat German as one language among dozens. Aleph Alpha Kolibri supports two, by design. The model card calls it “a deliberate choice of depth over breadth”, and the blog says the result is “bilingual by design, not an English model that has read some German”.

UniBPE and compound words

German builds long compound nouns, and standard tokenizers chop them at odd points. Aleph Alpha trained a new tokenizer method it calls UniBPE, a mix of the two common approaches, BPE and Unigram. Its example is “Bundessozialgerichtes” (of the Federal Social Court). Aleph Alpha Kolibri splits it as Bundes, sozial, gericht, es. GPT-5, Qwen3.8 and Gemini all split it as Bund, ess, oz, ial, gericht, es, which cuts “Bundes” in the wrong place.

TokenizerVocabularyGerman web, bytes per tokenEnglish web, bytes per token
Kolibri128,0004.904.58
GPT-5200,0194.354.67
Qwen3.5 to Qwen3.8248,0774.174.47
Gemini262,1444.134.49
Tekken (Mistral, Nemotron, Apertus)131,0724.034.45
GLM 5.3154,8563.934.61
DeepSeek V4129,2803.724.59
Kimi K3163,5863.284.62

More bytes per token means fewer tokens for the same text. Aleph Alpha’s tech report puts the German saving at 11.2% fewer tokens than GPT-5, the best of the nine other tokenizers it tested. The arithmetic checks out: 4.35 ÷ 4.90 = 0.888, so the same German text needs about 11% fewer tokens.

Finding four trillion German tokens

Small test models told Aleph Alpha that about 20% German was the best mix, which meant finding roughly 4 trillion German tokens for a 20-trillion-token run. After filtering, open German datasets supplied only 390 billion. Aleph Alpha built its own German Common Crawl pipeline, retuned so it no longer threw out long administrative words, which yielded 1.3 trillion tokens. Rewriting German documents in new styles added about a trillion more. Kolibri saw the resulting 2.4-trillion-token pool about 1.8 times on average.

Why it matters for cost

Fewer tokens per document means lower bills, shorter waits and more room in the context window. A German public body sending long legal texts through Aleph Alpha Kolibri pays for about 11% fewer tokens than with a GPT-5-class tokenizer, before any difference in model price. The model also reasons in German rather than only answering in it, which makes its reasoning easier for German-speaking staff to check.

Running Aleph Alpha Kolibri on Your Own Hardware

The default download is an FP8 checkpoint of about 78 GB. A separate BF16 repository holds full-precision weights at about 156 GB. Here is what Aleph Alpha says each needs.

VersionWeights in memoryMinimumRecommended
Kolibri-1 (FP8)About 78 GB2x A100 80 GB, 2x H100 SXM5, 1x H200, 1x B200 or 1x B3002x H100 SXM5, 2x H200, 1x B200 or 1x B300
Kolibri-1-BF16About 156 GB4x A100 80 GB, 4x H100 SXM5, 2x H200, 1x B200 or 1x B3004x H100 SXM5, 2x H200, 2x B200 or 1x B300

Those are data-centre GPUs. The 3.46 billion active parameters make each token cheap to compute, but all 78 billion still have to sit in GPU memory. For smaller machines, community builds quantised to 2, 3, 4, 6 and 8 bits appeared on Hugging Face within two days, though Aleph Alpha has not tested them. Our local LLM hardware guide covers what consumer cards can realistically hold.

The vLLM plugin

Aleph Alpha Kolibri does not run on stock vLLM. You need Aleph Alpha’s aleph-alpha-inference package, which adds the model code plus Kolibri-specific parsers for reasoning and tool calls. It is Apache 2.0 too, and also ships as a container image. Its package file pins vLLM to the 0.29 series, with the note that “each release supports exactly one vLLM minor version”. Expect to wait for a plugin update before moving to a newer vLLM.

Settings that matter

The model card recommends a temperature of 1.0, top_p of 0.97 and top_k of 128. Reasoning effort is set per request through the chat template: none, low, medium or high. Higher effort means more thinking tokens, so it costs more and answers more slowly. The server exposes an OpenAI-compatible API, so existing client code works by changing the base URL and model name.

Where it fits among open-weight options

Aleph Alpha Kolibri sits at the small, cheap-to-serve end of the field. A much larger mixture-of-experts model such as DeepSeek V4.1 Flash at 552B total needs far more memory. Our guide to open-weight AI models in 2026 compares the wider field.

The Licence: What Apache 2.0 Covers for Aleph Alpha Kolibri

Apache 2.0 is one of the most permissive licences there is. You can use the weights commercially, modify them and redistribute them, with no user caps and no field-of-use limits.

Weights yes, recipe no

The model card adds a limit worth reading. The licence “only applies to the weights and configuration files published in this repository” and “does not extend to underlying code, model architecture, parameter settings or any training method”. Aleph Alpha keeps all rights to its training code and methods. So Aleph Alpha Kolibri is open-weight, not open-source in the fuller sense that projects such as OLMo use, where data and training code are also released.

Guidance, not conditions

The responsible-use section asks people to avoid prohibited practices under Article 5 of the EU AI Act and activities “related to military or nuclear applications”. Because the weights are under Apache 2.0, these are requests, not licence terms. One launch report, from Crypto Briefing, described the model as pitched at “public administration, industry, aerospace, and defense”. Aleph Alpha’s own blog names public administration, industrials and aerospace, and does not mention defence.

Where it sits under the EU AI Act

Aleph Alpha says it built Aleph Alpha Kolibri with the EU AI Act, the GPAI Code of Practice and GDPR in mind, and it is a Code of Practice signatory. Training used 6.4 × 10^23 floating-point operations, according to the model card. The Act presumes systemic risk above 10^25, so Kolibri used about 6.4% of that threshold. Under Article 53(2), open-source models below that line are excused two documentation duties, but still need a copyright policy and a public summary of training content. Aleph Alpha has published that summary on the European Commission’s template.

The energy bill

The model card estimates training energy at 950 MWh, including data-centre overhead. That covers pre-training, mid-training and long-context extension, which used 392,000, 90,000 and 10,000 GPU-hours, so 492,000 in total. It excludes fine-tuning and reinforcement learning. Dividing 950,000 kWh by 492,000 GPU-hours gives about 1.9 kWh per B200 GPU-hour, all overheads included.

Aleph Alpha Kolibri and the Cohere Merger

Aleph Alpha Kolibri lands at an odd moment for the company. Aleph Alpha is being absorbed by Cohere, the Toronto-based enterprise AI company, and the release may be one of the last under the Aleph Alpha name.

WhenWhat happened
April 2026Cohere and Aleph Alpha announce a planned combination, with Schwarz Group backing
11 June 2026Kolibri Origin finishes pre-training
11 September 2026Kolibri finishes pre-training
16 September 2026Definitive business combination agreement signed
3 October 2026Kolibri released under Apache 2.0
Later in 2026Deal expected to close, subject to regulatory approval

The deal

We covered the plan when it was announced in our piece on the Cohere Aleph Alpha merger. The definitive agreement, reported by The Next Web on 16 September, keeps the Cohere name, with headquarters in Berlin and Toronto and more than 1,000 staff. Heidelberg stays open as a research centre. Schwarz Group, owner of Lidl and Kaufland, plans to invest €500 million, and the combined company is reportedly valued at about $20 billion.

What it means for Aleph Alpha Kolibri

The weights are out under Apache 2.0, and that cannot be taken back. Anyone who downloads Aleph Alpha Kolibri today can keep using it whatever happens to the brand. What is less certain is the roadmap. Aleph Alpha co-founder Samuel Weinbach is due to become Cohere’s chief research officer and co-chief executive Ilhan Scheer its chief operating officer. Cohere sells its own Command models, and neither company has said whether a Kolibri 2 would ship under that name.

Who Should Look at Aleph Alpha Kolibri

Strong fit

Public bodies and regulated firms in German-speaking countries that need a model on their own servers are the obvious audience. So are teams building document search and question answering where a wrong answer costs more than no answer. And anyone who needs cheap long-context serving will find the sliding-window design useful, provided they stay near the 262,144-token native window.

Weaker fit

Teams that need the best open model for coding agents should look at Qwen3.6 or Qwen3.5 first, going by Aleph Alpha’s own table. Anyone working in French, Spanish or any language other than German and English is outside what Aleph Alpha Kolibri was built for. And if your workload depends on recalling facts from training rather than from supplied documents, its low accuracy on knowledge tests matters.

Test it before you trust the 1M figure

If the million-token window is the reason you are interested, run your own documents through it at the length you need. Aleph Alpha’s own numbers show the model works at that length, with a lower score, and its own advice is to stay at 262,144 tokens or below for demanding work. For most document workloads that is still a very large window.

Aleph Alpha Kolibri FAQ

What is Aleph Alpha Kolibri?

Aleph Alpha Kolibri is an open-weight English and German language model released on 3 October 2026. It is a mixture-of-experts model with 78.1 billion total parameters and 3.46 billion active per token, with reasoning and tool calling, under the Apache 2.0 licence.

Does Aleph Alpha Kolibri really have a 1M context window?

It can be served at 1,048,576 tokens, and Aleph Alpha has tested it there. It was trained only up to 262,144 tokens, which is its native length, and the model card recommends staying at or below that for demanding tasks. Its RULER score falls from 69.8 at 256K to 63.2 at 1M.

What hardware does Aleph Alpha Kolibri need?

The FP8 version needs about 78 GB of GPU memory: at minimum two A100 80 GB or two H100 cards, or one H200, B200 or B300. The BF16 version needs about 156 GB. Community quantisations are smaller but untested by Aleph Alpha.

Is Aleph Alpha Kolibri open source?

The weights are under Apache 2.0, so you can use, change and redistribute them commercially. The training code, data and methods are not released, so it is open-weight rather than fully open-source.

How does Aleph Alpha Kolibri compare with Qwen and Mistral?

In Aleph Alpha’s tests it beats Qwen3.6 35B-A3B, Qwen3.5 35B-A3B and Mistral Small 4 on overall English and German scores and on maths. It trails both Qwen models on coding agents and tool use. Newer models such as GLM-5.3, Kimi K3 and the larger Qwen3.8 releases were not in its main comparison.

What happens to Kolibri after the Cohere merger?

The released weights stay available under Apache 2.0 regardless. The combined company will operate as Cohere once regulators approve the deal, with Heidelberg kept as a research centre. Neither company has said what the next Kolibri release will be called.

References and Further Reading