Aleph Alpha launches Kolibri open model with 78B parameters, 3B active and 1M-token context

Aleph Alpha has released Kolibri, a new German-English open-weight Mixture-of-Experts model with 78.1 billion total parameters, 3.46 billion active parameters per token and a maximum context window of 1,048,576 tokens.
The model combines controllable reasoning, tool calling and long-context processing with an architecture designed to keep inference requirements substantially below what its 78B headline parameter count might suggest.
Kolibri is available with downloadable weights under the Apache 2.0 license. Aleph Alpha positions it for multi-step reasoning, retrieval-augmented generation, agentic tool use, coding and German- and English-language enterprise assistants.
The release also represents a major scaling step from the company's unreleased Kolibri Origin prototype: total parameters increase from 30.6B to 78.1B and maximum supported context expands from 65,536 to more than one million tokens, while the number of parameters activated for each token remains close to 3B.
··········
KOLIBRI AT A GLANCE
........
Specification | Kolibri |
Developer | Aleph Alpha |
Architecture | Mixture-of-Experts |
Total parameters | 78.1B |
Active parameters per token | 3.46B |
Languages | German and English |
Maximum context | 1,048,576 tokens |
Recommended context | Up to 262,144 tokens |
Reasoning | None, Low, Medium, High |
Tool calling | Yes |
Number of layers | 50 |
Experts | 384 total / 6 active |
Training tokens | 20T pre-training |
Training hardware | 768 NVIDIA B200 GPUs |
Knowledge cutoff | June 18, 2026 |
Model weights | Open |
License | Apache 2.0 |
........
The distinction between total and active parameters is central to Kolibri's design.
The complete model contains more than 78 billion parameters, but only approximately 3.46 billion are activated for each token. This allows the model to maintain a much larger collection of specialized parameters without incurring the compute cost of executing the entire network for every token.
··········
KOLIBRI USES 384 EXPERTS BUT ACTIVATES ONLY SIX PER TOKEN
Kolibri uses a highly sparse Mixture-of-Experts architecture.
Its 50 transformer layers are all MoE layers. Each layer contains specialized feed-forward experts, with the routing system selecting a small subset according to the token being processed.
Kolibri contains 384 experts and activates six for each token, compared with 128 experts and eight active experts in the earlier Kolibri Origin design.
Increasing the total number of experts expands the model's parameter capacity while reducing the proportion of the network that needs to participate in an individual forward pass.
This produces the apparently unusual combination of 78.1B total parameters and only 3.46B active parameters per token.
The architecture does not make Kolibri computationally equivalent to a conventional dense 3.46B model. Attention, routing, memory traffic and the storage of all 78B parameters remain relevant to deployment. The active-parameter figure instead describes the amount of expert capacity selected during token processing.
··········
THE CONTEXT WINDOW REACHES 1,048,576 TOKENS
Kolibri supports a maximum context length of 1,048,576 tokens, or slightly more than one million tokens.
Aleph Alpha recommends operating at 262,144 tokens or below for efficient serving and complex workloads.
That distinction is important because maximum context capacity and practical context efficiency are not equivalent.
Longer sequences increase KV-cache requirements and the computational burden of attention. A model capable of accepting one million tokens can therefore be technically compatible with that context length without making it the optimal setting for every deployment.
Aleph Alpha trained Kolibri progressively for longer contexts.
Initial pre-training used sequences of 16,384 tokens. Mid-training expanded this to 65,536 tokens, followed by long-context adaptation at 262,144 tokens.
The final model can then be configured for inference at the full 1,048,576-token limit.
··········
ONLY ONE IN FIVE LAYERS USES FULL ATTENTION
Aleph Alpha also changed the attention architecture to reduce the cost of processing long sequences.
Of Kolibri's 50 layers, 10 use full attention, while the remaining 40 operate with a sliding attention window of 512 tokens.
Full attention allows tokens to interact across the entire available context but becomes increasingly expensive as sequence length grows.
Sliding-window attention restricts most layers to a much smaller local region.
Kolibri therefore alternates between inexpensive local processing and periodic layers capable of integrating information across the complete sequence.
........
Attention component | Kolibri configuration |
Total transformer layers | 50 |
Full-attention layers | 10 |
Sliding-window layers | 40 |
Sliding window | 512 tokens |
Full-attention frequency | Every fifth layer |
Maximum supported context | 1,048,576 tokens |
........
Data Studios calculation: 80% of Kolibri's transformer layers use the bounded 512-token sliding window, while 20% perform full-context attention.
This architecture is particularly relevant at long context lengths because the computational requirements of the sliding-window layers remain bounded even as the overall prompt becomes much larger.
··········
ALEPH ALPHA TRAINED KOLIBRI ON 20 TRILLION TOKENS
Kolibri's main pre-training stage used 20 trillion tokens.
Aleph Alpha says its data pipeline processed more than 200 trillion raw tokens before filtering, deduplication and curation reduced the dataset to the material ultimately used for training.
The model was trained on 768 NVIDIA B200 GPUs.
The main 20T-token pre-training phase ran for 21 days at a 16K sequence length. Aleph Alpha subsequently performed mid-training using 3.44 trillion tokens at 64K context and a further long-context adaptation phase using approximately 200 billion tokens at 256K.
The resulting knowledge cutoff is June 18, 2026 for both German and English.
The company emphasizes German data as a major part of the model rather than treating German primarily as a translated extension of an English-first system.
That makes Kolibri particularly relevant for European enterprise and government deployments where language quality, local infrastructure and control over model weights can be operational requirements rather than secondary features.
··········
KOLIBRI ADDS FOUR LEVELS OF REASONING EFFORT
Kolibri supports explicit reasoning with four operating levels:
None, Low, Medium and High.
This allows applications to vary the amount of reasoning performed according to the complexity of a request.
Routine extraction or classification tasks can avoid unnecessary reasoning overhead, while difficult analytical tasks can allocate additional inference to the reasoning process.
The model also supports native tool calling, making the reasoning controls relevant to agentic workflows.
An enterprise agent could therefore use lower reasoning effort for routine tool selection and increase reasoning effort when a task requires multi-stage planning, analysis or decisions involving several retrieved sources.
This creates a compute-quality control at inference time rather than forcing every request through the same reasoning profile.
··········
ALEPH ALPHA EXPANDED FROM 30B TO 78B WITHOUT SIGNIFICANTLY INCREASING ACTIVE PARAMETERS
Kolibri originated from an earlier internal model called Kolibri Origin.
The progression between the two illustrates the effect of sparse scaling.
........
Specification | Kolibri Origin | Kolibri |
Total parameters | 30.6B | 78.1B |
Active parameters/token | 3.27B | 3.46B |
Experts | 128 | 384 |
Active experts | 8 | 6 |
Training tokens | 7.51T | 20T |
Longest trained context | 65,536 | 262,144 |
Maximum supported context | 65,536 | 1,048,576 |
Reasoning settings | One | Four |
Vocabulary | 96,000 | 128,000 |
........
Data Studios calculation: total parameter count increased by approximately 155%, from 30.6B to 78.1B, while active parameters per token increased by only about 5.8%, from 3.27B to 3.46B.
Training-token volume increased approximately 166%, from 7.51T to 20T.
This is the central scaling strategy behind Kolibri: substantially increase total model capacity while keeping per-token expert activation nearly flat.
··········
ALEPH ALPHA REJECTED A LARGER 123B DESIGN FOR SERVING EFFICIENCY
Aleph Alpha experimented with models extending to 123 billion parameters and found that larger configurations continued to improve model performance.
It nevertheless selected the 78B architecture because the larger alternative imposed significantly higher serving costs.
According to Aleph Alpha's internal testing, a 123B configuration could process only three concurrent 256K-context requests on two H100 GPUs, while the selected 78B design handled 18 concurrent requests under the company's comparison.
Kolibri also decoded 28% faster than the larger configuration.
That represents a deliberate engineering trade-off between maximum benchmark capability and deployability.
For enterprise inference, model quality is only one variable. Concurrency, latency, memory requirements and infrastructure cost determine how many real users or agents can operate on the same hardware.
The final architecture therefore reflects a serving constraint rather than simply selecting the largest model produced during scaling experiments.
··········
KOLIBRI TARGETS ON-PREMISE ENTERPRISE AND GOVERNMENT AI
Aleph Alpha is positioning Kolibri around environments where organizations want direct control over inference infrastructure and model weights.
The official FP8 release requires approximately 78GB for model weights.
Aleph Alpha lists configurations including a single H200, B200 or B300 among the minimum deployment options, while two A100 80GB or two H100 SXM5 GPUs can also run the model.
The company recommends stronger configurations for production workloads, particularly when long contexts and concurrency are required.
Open weights allow organizations to deploy Kolibri without sending prompts, retrieved documents and internal data to an external inference provider.
This is particularly relevant for retrieval systems working with confidential corporate records, government documents, regulated information or intellectual property.
Open deployment does not by itself provide security or regulatory compliance, but it gives organizations direct control over the infrastructure on which those requirements must be implemented.
··········
ALEPH ALPHA REPORTS STRONG RESULTS FOR ITS ACTIVE-PARAMETER CLASS
Aleph Alpha's own benchmark suite reports an overall English score of 75.5 and German score of 70.8 for Kolibri.
The company compares the model with several open and open-weight systems including Qwen, GLM, Nemotron, Gemma, GPT-OSS, Mistral and Apertus.
On GPQA Diamond, Aleph Alpha reports 84.3 in English and 81.3 in German.
The results are particularly notable relative to the model's approximately 3.46B active parameters per token, although total memory requirements remain determined by the much larger 78B parameter pool.
These figures are Aleph Alpha's own evaluations, not independent benchmark results, and should therefore be interpreted as vendor-reported measurements until broader third-party testing becomes available.
The release is new enough that independent analysis of real-world throughput, long-context reliability, tool use and reasoning behavior remains limited.
··········
OPEN WEIGHTS MAKE KOLIBRI MODIFIABLE RATHER THAN API-ONLY
Kolibri is distributed under the Apache 2.0 license, allowing developers and organizations to download and deploy the model weights.
Aleph Alpha provides an inference package integrating the model with vLLM, including parsers for Kolibri's reasoning output and tool calls.
The model is therefore not restricted to an Aleph Alpha-hosted API.
Organizations can control serving infrastructure, integrate the model into private retrieval systems and build specialized applications around its weights.
That distinction is increasingly important as enterprise AI separates into two deployment models: remotely hosted frontier APIs optimized for maximum general capability, and open-weight systems optimized for control, customization and infrastructure ownership.
Kolibri is explicitly designed for the second category.
··········
KOLIBRI COMBINES SPARSE SCALE WITH EUROPEAN AI SOVEREIGNTY
Kolibri's headline specification is 78B parameters, but the model's architecture is built around avoiding the cost normally associated with executing a dense model of that size.
Only 3.46B parameters are active per token. Most transformer layers use bounded sliding-window attention. Reasoning effort can be adjusted according to task complexity. The model can run on privately controlled infrastructure and its weights are available under Apache 2.0.
At the same time, Aleph Alpha has expanded training from 7.51T tokens in Kolibri Origin to 20T, tripled the number of experts and increased supported context from 65K to more than one million tokens.
The resulting model is not an attempt to maximize a single dimension such as parameter count or context length.
It combines sparse computation, long-context processing, controllable reasoning, tool calling and open deployment around an architecture designed for enterprise inference.
For Aleph Alpha, Kolibri also marks a return to the open-model competition with a distinctly European positioning: German and English capabilities, downloadable weights and infrastructure ownership as core characteristics rather than optional deployment features.
··········
FOLLOW US FOR MORE.
DATA STUDIOS
datastudios.org




