Privacy and compliance · 2026-10-04

Does Kolibri support Turkish? Sovereign open models for teams in Türkiye

Aleph Alpha Kolibri was trained for German and English only. See what that means for Turkish work, how the tokenizer affects it, and what open weights add to data residency.

LLMTR diagram comparing Kolibri's model-card support status and tokenizer efficiency for German, English and Turkish, prepared as an explanatory illustration.

Short answer: Turkish is not officially supported

Aleph Alpha writes that it trained and evaluated Kolibri to process German and English text. About 62.5 percent of the pre-training data is English, 23.9 percent German and 13.6 percent code. The company describes going deep in two languages as a deliberate choice over spreading across many. Turkish is not one of those languages, and the model card has no evaluation result for Turkish.

This does not mean the model cannot read Turkish at all; models trained on broad web data can produce some text in unsupported languages. But that behaviour has been measured neither by the vendor nor by LLMTR. Any judgement about Kolibri's Turkish quality should rest on a measurement you run on your own examples.

Measuring does not need a large dataset. Prepare a few dozen examples taken from your real work: typical questions, expected answers and a simple right-or-wrong criterion for each answer. Run the same set on Kolibri and on a model that has been evaluated on Turkish. Count grammar errors, answers that drift into German or English, and the consistency of domain terms separately; a single overall score hides these differences.

Language support and the tokenizer

Kolibri's tokenizer was trained with a 128,000-entry vocabulary and a method called UniBPE that respects German word structure, compound words in particular. The more bytes per token, the fewer tokens the same text is split into; that affects both how much text fits in context and the cost on a service billed per token.

Turkish also builds long words by adding suffixes, but how efficiently the tokenizer splits Turkish text has not been published. Measure instead of guessing: split a sample of your own Turkish documents with the tokenizer from the Hugging Face repository and divide the text's byte count by its token count. Measure the same sample on the other models you use and the comparison becomes meaningful.

Kolibri-1 language support and token efficiency, Aleph Alpha model card as of 3 October 2026
LanguageStatus in the model cardBytes/token (Aleph Alpha)
GermanTrained and evaluated4.7
EnglishTrained and evaluated4.2
TurkishNot documentedNot published

What do open weights add to data residency?

In organisations, two questions usually get mixed up in the AI debate: where data is processed, and whom the model depends on. Kolibri gives a strong answer to the second. Because the weights are open under Apache 2.0, the model can run on chosen infrastructure without being tied to a single provider's service; if a provider changes its price or terms, you keep a portable option.

The first question is answered not by the model but by the infrastructure that runs it. If a model trained in Europe is hosted in another country, requests are processed in that country. For KVKK and public-sector rules, what decides the matter is not where the model comes from but where, and through which chain of sub-processors, the request is processed. The sovereign model label does not answer that question for you.

A checklist for enterprise use

When evaluating Kolibri for an internal project, answer these five points separately:

  • Task language: if the work is in German or English it falls within the model card's evaluations; if it is in Turkish, do not ship it without measuring on your own evaluation set.
  • Data residency: confirm separately, with the infrastructure that serves it, which country the model runs in.
  • Licence: Apache 2.0 permits commercial use and covers the files published in the repository.
  • Human oversight: Aleph Alpha recommends the model for systems in which a person reviews the output before it is acted on.
  • Knowledge cutoff: the model's knowledge ends on 18 June 2026; newer information needs tool calls or attached documents.

Where to look on LLMTR for Turkish-heavy work

If the quality of Turkish text is decisive, starting with models that have been evaluated on Turkish separately is the safer path. In the LLMTR catalog you can compare models by context length and price, and models hosted in Türkiye or the EU carry their own marker; you can make the comparison on the /models page.

Kolibri-1 is available free on the LLMTR API as tesseracted/kolibri-1 until 9 October 2026, 23:59 (TRT); after that date access closes and the model page stays open. For German or English work you can compare Kolibri with existing models on the same evaluation set during that period.

Frequently asked questions

Can Kolibri produce Turkish text?

The model card does not list Turkish among the supported languages and publishes no evaluation result for it. The quality of any text it produces has not been measured; test with your own examples before using it for Turkish work.

Does an open-weight model provide KVKK compliance on its own?

No. Open weights do not decide where the model runs. For data protection, what decides the matter is the infrastructure that processes the request and that infrastructure's chain of sub-processors.

Which region does Kolibri run in on LLMTR?

Kolibri-1 on LLMTR runs on Tesseracted Labs infrastructure; the provider has not published a serving region, so the model page carries no region label. It is not one of the models LLMTR hosts in Türkiye. Under the provider's terms, prompt and response content is not stored.

Related posts