Summary
- Kolibri uses a mixture of experts architecture with about 78 billion total parameters and roughly 3.5 billion active for each token.
- The model supports German and English, contexts up to one million tokens and deployment on infrastructure controlled by customers.
- Aleph Alpha has released the model weights under the Apache 2.0 licence as European suppliers compete around sovereignty and enterprise deployment.
Aleph Alpha has released the weights of its new Kolibri language model under the Apache 2.0 licence, combining a large European trained model with the option for organisations to operate it on infrastructure they control.
Aleph Alpha describes Kolibri as a German and English mixture of experts model with around 78 billion parameters in total and approximately 3.5 billion active for each token processed. Its weights are available through Hugging Face, while customers can deploy the model through Aleph Alpha’s inference software rather than being restricted to a centrally hosted API.
The model supports reasoning and tool calling and can handle context lengths of up to one million tokens, although the company recommends 262,144 tokens for more efficient operation. German material represents around 23% of its pretraining data, reflecting Aleph Alpha’s attempt to optimise the system for a language that receives less attention from the largest US model developers.
Making the weights available gives technical substance to an increasingly broad sovereignty claim because organisations can download a particular model version, choose where it is hosted and retain greater control over the infrastructure processing their information.
Sovereignty becomes an architecture choice
European AI sovereignty is often discussed through the location of data centres or nationality of suppliers, but operating control depends on several layers at once. Infrastructure location, model access, software dependencies, update mechanisms and contractual authority can all determine whether an organisation remains dependent on an external provider.
Openly available model weights change part of that relationship because an enterprise or public authority can retain a specific version and run it within an environment it manages itself. That can be useful where sensitive information, network isolation or procurement requirements make a conventional public AI service unsuitable.
Local control does not create complete technological independence because Kolibri still requires substantial accelerator hardware and supporting software, while the infrastructure beneath a deployment may depend on non-European semiconductor and cloud suppliers. Customers also need engineering capability to operate, monitor and update the system.
Even so, access to the weights gives organisations more architectural choice than a service available only through a provider’s API and makes the model easier for researchers and developers to inspect, benchmark and adapt.
Kolibri arrives shortly after Aleph Alpha and Cohere agreed their combination, a deal intended to build greater scale around AI for regulated and enterprise customers. Releasing a substantial model under a permissive licence indicates that the enlarged strategy will not depend solely on proprietary hosted access.
Efficiency determines whether control is affordable
Local deployment becomes commercially useful only when the hardware requirement remains manageable. Kolibri’s mixture of experts architecture is designed to reduce the amount of the model used during any individual inference operation rather than activating all 78 billion parameters for every token.
Aleph Alpha says roughly 3.46 billion parameters are active for each token. The architecture contains 384 experts, of which six are selected during processing, allowing the model to retain a much larger total parameter pool without using the entire network simultaneously.
The company lists configurations ranging from two Nvidia A100 80GB GPUs to individual newer B200 or B300 accelerators among supported hardware options, depending on deployment. Those requirements still place serious local operation beyond an ordinary business server but bring a large model within reach of enterprise and government compute environments.
The same economics affect context length because maximum technical capability and sensible production configuration are not necessarily the same thing. Although the architecture can support one million tokens, Aleph Alpha recommends a shorter 262,144-token operating context for efficient complex tasks.
Enterprises comparing models by parameter counts and context windows therefore still have to consider cost per useful task, latency, hardware availability and operational complexity before deciding whether a system can be deployed widely.
German optimisation is part of the proposition
Kolibri was trained on around 20 trillion tokens and uses a tokenizer designed to represent German text more efficiently. Language specific optimisation can affect cost as well as quality because inefficient tokenisation forces a model to use more tokens to represent the same passage.
A supplier targeting German public administration and industry can therefore compete on factors that receive less attention in general English language benchmarks. Legal documents, technical terminology, compound words and administrative language create demands that can disappear inside aggregate multilingual scores.
Aleph Alpha says the model was developed and trained in Europe and has documented data provenance and design decisions. Those characteristics are aimed particularly at public and regulated organisations that need to understand more about the systems they procure than whether they lead a general reasoning benchmark.
The open licence also changes part of the procurement equation because customers can evaluate the same weights themselves and potentially build internal expertise around them, reducing some of the switching risk associated with a purely hosted model.
Kolibri will still have to win production workloads against increasingly capable global competitors on accuracy, reliability, operating cost and ecosystem support rather than sovereignty claims alone.
Its release nevertheless makes the sovereignty argument more concrete. Instead of asking organisations simply to trust that a model is operated in an acceptable jurisdiction, Aleph Alpha is giving them the option to take the model itself and decide where it runs.












