Model comparisons · 2026-10-04
What is Aleph Alpha Kolibri? An open-weight model trained in Germany
The architecture, context length, language support and licence of Kolibri, released by Aleph Alpha under Apache 2.0 on 3 October 2026, and its status on LLMTR.
Short answer: what is Kolibri?
Kolibri is a large language model released on 3 October 2026 by Aleph Alpha, a company based in Heidelberg. It was trained for German and English, and its weights were published on Hugging Face as Aleph-Alpha/Kolibri-1 under the Apache 2.0 licence, which also permits commercial use. Aleph Alpha presents it as a sovereign, open-weight model: built in Germany and trained on infrastructure in Germany and Finland under German and European law.
Kolibri is distributed as weights, not as a chat app or a public API. To use it you either run it on your own servers or call it from a platform that serves it. Aleph Alpha's target audience follows from that: public bodies that want to keep data on their own infrastructure, and sectors with heavy compliance requirements such as industry, aerospace and automotive.
Technical specifications
The values below are taken from Aleph Alpha's model card. In Aleph Alpha's own evaluation Kolibri scores 75.5 overall in English and 70.8 overall in German. These results are the vendor's statement; LLMTR did not measure them.
| Field | Kolibri-1 |
|---|---|
| Total parameters | 78.1 billion |
| Active parameters per token | 3.46 billion |
| Architecture | Mixture-of-Experts, 50 layers, 384 experts + 1 shared expert per layer, 6 experts per token |
| Context | 262,144 tokens trained length, validated up to 1,048,576 tokens |
| Languages | German, English |
| Input and output | Text only |
| Reasoning levels | none, low, medium, high |
| Tool calling | Yes |
| Knowledge cutoff | 18 June 2026 |
| Licence | Apache 2.0 |
78 billion parameters, 3.46 billion active: what does MoE mean?
In a Mixture-of-Experts architecture every layer of the model holds many small sub-networks, the experts. For each token a router picks only a few of them. In Kolibri 6 of the 384 experts per layer run, plus one shared expert, so only 3.46 billion of the 78.1 billion total parameters take part in computing each token.
The practical effect cuts both ways. Compute per token is close to that of a small model, which lowers latency and serving cost. Memory, on the other hand, does not shrink: because it is not known in advance which experts will be picked, the whole model has to sit in memory. Aleph Alpha states this trade-off openly in the model card and gives a footprint of about 78 GB with FP8 weights.
The Apache 2.0 licence and the sovereign model label
Aleph Alpha is not new to this field. The company earlier had the Luminous model family, and the weights of its 7-billion-parameter Pharia-1-LLM models released in 2024 were also available, but the Open Aleph License only allowed non-commercial research and educational use. The real difference with Kolibri is the licence: Apache 2.0 allows you to download the model, modify it and run it inside a commercial product.
The sovereign model label is Aleph Alpha's own positioning and rests on three things: the model was developed and trained in Europe, the training chain is documented from data collection to final evaluation, and open weights let the model run on infrastructure the user chooses. Open weights alone do not decide where data is processed; that is a property of the infrastructure serving the model.
What is it designed for?
Aleph Alpha recommends Kolibri for multi-step reasoning, question answering over documents, agent workflows that call tools, coding, and German and English assistants. The model card also says it is designed for systems in which a person reviews the output before it is acted on, and that it belongs on the side that prepares options rather than the component that makes the decision.
Long context matters for these uses. The model was trained with a 262,144-token context and Aleph Alpha validated quality up to 1,048,576 tokens; even so, it recommends staying at 262,144 tokens for latency-sensitive work and complex tasks. The model's knowledge ends on 18 June 2026; newer information needs tool calls or attached documents.
Status on LLMTR
Kolibri-1 is available free on the LLMTR API as tesseracted/kolibri-1 until 9 October 2026, 23:59 (TRT); after that date access closes and the model page stays open.
The table above describes the model itself. The offering on LLMTR has different limits: input and output together must fit in 32,768 tokens and one response returns at most 16,384 tokens. Reasoning is off by default and is turned on with reasoning_effort; tool calling is supported. The first request and the full limits are on the Kolibri-1 model page and in the guide to using it on LLMTR.
Frequently asked questions
Does Kolibri support Turkish?
Aleph Alpha says it trained and evaluated the model for German and English only; Turkish is not among the officially supported languages. LLMTR has not measured the model's Turkish performance.
Does Kolibri have a public API?
As of 3 October 2026 Aleph Alpha has published the model as weights on Hugging Face and points to its own team for enterprise deployment; it has not announced a pay-as-you-go API. You can run the model on your own servers with vLLM, or try it free on the LLMTR API until 9 October 2026, 23:59 (TRT).
Is Kolibri available on LLMTR?
Yes. It is available free as tesseracted/kolibri-1 until 9 October 2026, 23:59 (TRT). After that date access closes and the model page stays open.
Related posts
- Running Kolibri on your own servers: hardware, vLLM and context settings
- Does Kolibri support Turkish? Sovereign open models for teams in Türkiye
- AI Sovereignty and Open-Weight Models: What It Means for Europe
- Kolibri-1 API: free on LLMTR until 9 October, first request and limits
- Kolibri-1 tool calling and reasoning_effort: a developer guide for the LLMTR API