Integration guides · 2026-10-04

Kolibri-1 API: free on LLMTR until 9 October, first request and limits

Aleph Alpha's open-weight model Kolibri-1 is free on the LLMTR API until 9 October 2026, 23:59 (TRT), with the support of Tesseracted Labs. Model id, first request, limits and what follows.

LLMTR diagram showing the model id, the request endpoint, the context limit and the last day of the free period for Kolibri-1.

Short answer: free until 9 October 2026, 23:59

Kolibri-1, which Aleph Alpha released under Apache 2.0 on 3 October 2026, is available on the LLMTR API as tesseracted/kolibri-1 from 4 October 2026. The model is free: requests are not charged to your balance and you do not need to top up to use it. The free period ends on 9 October 2026 at 23:59 (Türkiye time).

Kolibri-1 is a model of European origin: it was developed by Aleph Alpha, based in Heidelberg, and trained for German and English. Aleph Alpha recommends it for multi-step reasoning, agent flows that call tools, document-grounded question answering and coding. We described the model's architecture and licence in a separate article; this one covers its use on LLMTR: the first request, the limits and what happens when the free period ends.

Thanks to Tesseracted Labs

Tesseracted Labs is the reason we can offer Kolibri-1 to LLMTR users at no charge. The Germany-based team agreed to open the model for LLMTR for free; the free period in this article is the result of that support. We thank Tesseracted Labs for it.

Send the first request: curl

Create an LLMTR API key and send the request to https://llmtr.com/v1/chat/completions. Write the id in the model field in full, with its provider prefix: tesseracted/kolibri-1. The example below reads the key from the LLMTR_API_KEY environment variable.

The example prompt is in German because the model was trained for German and English. If you do not send max_tokens the response is capped at 4,096 tokens; a smaller value is enough for a short answer. The usage field of the response shows prompt_tokens and completion_tokens.

POST /v1/chat/completions: first request with curl

curl https://llmtr.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMTR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "tesseracted/kolibri-1",
  "messages": [
    {
      "role": "system",
      "content": "Antworte kurz und sachlich."
    },
    {
      "role": "user",
      "content": "Erkläre in zwei Sätzen, was ein Mixture-of-Experts-Modell ist."
    }
  ],
  "max_tokens": 400
}'

The same request with the OpenAI SDK: Python

If you use the OpenAI SDK, changing two things is enough: the base URL becomes https://llmtr.com/v1 and the key becomes your LLMTR key. The rest of the code stays the same.

Kolibri-1 accepts only the fields it knows. If you send fields such as seed, frequency_penalty or user, the request is not rejected: those fields are not passed to the model and are listed by name in the llmtr_dropped_parameters field of the response. That list tells you whether your client relies on one of them.

/v1/chat/completions request with the OpenAI Python SDK

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://llmtr.com/v1",
    api_key=os.environ["LLMTR_API_KEY"],
)

response = client.chat.completions.create(
    model="tesseracted/kolibri-1",
    messages=[
        {"role": "system", "content": "Antworte kurz und sachlich."},
        {"role": "user", "content": "Erkläre in zwei Sätzen, was ein Mixture-of-Experts-Modell ist."},
    ],
    max_tokens=400,
)

print(response.choices[0].message.content)
print(response.usage)

Limits

The 262,144-token context on the model card belongs to the model itself. In this offering on LLMTR, input and output together must fit in 32,768 tokens and one response returns at most 16,384 tokens. If the input plus max_tokens exceeds 32,768, the request gets a 400; the error message says how many tokens the input uses and how much room is left for max_tokens.

The conversation history is limited to 80 messages and 60,000 characters. The model accepts text only; image input, JSON mode and JSON Schema output are not supported. Requests that ask for JSON through response_format get a 400. If you need structured output, describe the format in the prompt and validate the response on your side.

Limits of the tesseracted/kolibri-1 card
FeatureValue
Context (input and output together)32,768 tokens
Maximum output per response16,384 tokens
When max_tokens is omitted4,096 tokens
Conversation historyAt most 80 messages and 60,000 characters
InputText only
Tool callingSupported
JSON mode and JSON Schema outputNot supported
reasoning_effortnone (default), low, medium, high
End of the free period9 October 2026, 23:59 (TRT)

Language, data and error responses

Aleph Alpha says it trained and evaluated the model for German and English; Turkish is not among those languages. Test it on your own examples before using it for Turkish work. The model runs on Tesseracted Labs infrastructure; the provider has not published a serving region, and it is not one of the models LLMTR hosts in Türkiye. Under the provider's terms, prompt and response content is not stored; usage records such as token counts are kept for about 31 days.

Free models on LLMTR come with a daily usage quota. When the quota is used up you get a 429 and can try again the next day. If you send many requests in a short time you may also see a temporary 429; use exponential backoff in your client. At busy times the model may not be able to take a new request straight away: the request is not held waiting, it returns 503 model_busy with a Retry-After header. Wait for the time in the header and repeat the request.

What happens after 9 October

From 10 October 2026, requests sent to tesseracted/kolibri-1 get a 410 model_retired response. There is no other Aleph Alpha model in the LLMTR catalog, so the response does not suggest a replacement and no request is routed to another model on its own. Nothing is charged to your account.

The model page stays open and shows that the model is retired. If you want to keep a flow you built on Kolibri-1 running, pick another model on the /models page before 9 October and try the same prompts with it; apart from the model id the shape of the request does not change. To try it without writing code, you can also pick Kolibri-1 in the model selector of LLMTR Chat during the free period.

Frequently asked questions

Do I need to top up to use Kolibri-1?

No. An LLMTR account and an API key are enough. Requests are not charged to your balance; the model is offered free with a daily usage quota.

Why is Kolibri-1 limited to 32,768 tokens when the model card says 262,144?

262,144 tokens is the context length the model was trained with. In the free offering on LLMTR, input and output together must fit in 32,768 tokens; one response returns at most 16,384 tokens.

What happens when the free period ends?

After 9 October 2026, 23:59 (TRT), requests get a 410 model_retired response. No replacement is suggested because the catalog has no other Aleph Alpha model; the model page stays open.

Related posts