Integration guides · 2026-10-04
Kolibri-1 API: free on LLMTR until 9 October, first request and limits
Aleph Alpha's open-weight model Kolibri-1 is free on the LLMTR API until 9 October 2026, 23:59 (TRT), with the support of Tesseracted Labs. Model id, first request, limits and what follows.
Short answer: free until 9 October 2026, 23:59
Kolibri-1, which Aleph Alpha released under Apache 2.0 on 3 October 2026, is available on the LLMTR API as tesseracted/kolibri-1 from 4 October 2026. The model is free: requests are not charged to your balance and you do not need to top up to use it. The free period ends on 9 October 2026 at 23:59 (Türkiye time).
Kolibri-1 is a model of European origin: it was developed by Aleph Alpha, based in Heidelberg, and trained for German and English. Aleph Alpha recommends it for multi-step reasoning, agent flows that call tools, document-grounded question answering and coding. We described the model's architecture and licence in a separate article; this one covers its use on LLMTR: the first request, the limits and what happens when the free period ends.
Thanks to Tesseracted Labs
Tesseracted Labs is the reason we can offer Kolibri-1 to LLMTR users at no charge. The Germany-based team agreed to open the model for LLMTR for free; the free period in this article is the result of that support. We thank Tesseracted Labs for it.
Send the first request: curl
Create an LLMTR API key and send the request to https://llmtr.com/v1/chat/completions. Write the id in the model field in full, with its provider prefix: tesseracted/kolibri-1. The example below reads the key from the LLMTR_API_KEY environment variable.
The example prompt is in German because the model was trained for German and English. If you do not send max_tokens the response is capped at 4,096 tokens; a smaller value is enough for a short answer. The usage field of the response shows prompt_tokens and completion_tokens.
POST /v1/chat/completions: first request with curl
curl https://llmtr.com/v1/chat/completions \
-H "Authorization: Bearer $LLMTR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tesseracted/kolibri-1",
"messages": [
{
"role": "system",
"content": "Antworte kurz und sachlich."
},
{
"role": "user",
"content": "Erkläre in zwei Sätzen, was ein Mixture-of-Experts-Modell ist."
}
],
"max_tokens": 400
}'
The same request with the OpenAI SDK: Python
If you use the OpenAI SDK, changing two things is enough: the base URL becomes https://llmtr.com/v1 and the key becomes your LLMTR key. The rest of the code stays the same.
Kolibri-1 accepts only the fields it knows. If you send fields such as seed, frequency_penalty or user, the request is not rejected: those fields are not passed to the model and are listed by name in the llmtr_dropped_parameters field of the response. That list tells you whether your client relies on one of them.
/v1/chat/completions request with the OpenAI Python SDK
import os
from openai import OpenAI
client = OpenAI(
base_url="https://llmtr.com/v1",
api_key=os.environ["LLMTR_API_KEY"],
)
response = client.chat.completions.create(
model="tesseracted/kolibri-1",
messages=[
{"role": "system", "content": "Antworte kurz und sachlich."},
{"role": "user", "content": "Erkläre in zwei Sätzen, was ein Mixture-of-Experts-Modell ist."},
],
max_tokens=400,
)
print(response.choices[0].message.content)
print(response.usage)
Limits
The 262,144-token context on the model card belongs to the model itself. In this offering on LLMTR, input and output together must fit in 32,768 tokens and one response returns at most 16,384 tokens. If the input plus max_tokens exceeds 32,768, the request gets a 400; the error message says how many tokens the input uses and how much room is left for max_tokens.
The conversation history is limited to 80 messages and 60,000 characters. The model accepts text only; image input, JSON mode and JSON Schema output are not supported. Requests that ask for JSON through response_format get a 400. If you need structured output, describe the format in the prompt and validate the response on your side.
| Feature | Value |
|---|---|
| Context (input and output together) | 32,768 tokens |
| Maximum output per response | 16,384 tokens |
| When max_tokens is omitted | 4,096 tokens |
| Conversation history | At most 80 messages and 60,000 characters |
| Input | Text only |
| Tool calling | Supported |
| JSON mode and JSON Schema output | Not supported |
| reasoning_effort | none (default), low, medium, high |
| End of the free period | 9 October 2026, 23:59 (TRT) |
Language, data and error responses
Aleph Alpha says it trained and evaluated the model for German and English; Turkish is not among those languages. Test it on your own examples before using it for Turkish work. The model runs on Tesseracted Labs infrastructure; the provider has not published a serving region, and it is not one of the models LLMTR hosts in Türkiye. Under the provider's terms, prompt and response content is not stored; usage records such as token counts are kept for about 31 days.
Free models on LLMTR come with a daily usage quota. When the quota is used up you get a 429 and can try again the next day. If you send many requests in a short time you may also see a temporary 429; use exponential backoff in your client. At busy times the model may not be able to take a new request straight away: the request is not held waiting, it returns 503 model_busy with a Retry-After header. Wait for the time in the header and repeat the request.
What happens after 9 October
From 10 October 2026, requests sent to tesseracted/kolibri-1 get a 410 model_retired response. There is no other Aleph Alpha model in the LLMTR catalog, so the response does not suggest a replacement and no request is routed to another model on its own. Nothing is charged to your account.
The model page stays open and shows that the model is retired. If you want to keep a flow you built on Kolibri-1 running, pick another model on the /models page before 9 October and try the same prompts with it; apart from the model id the shape of the request does not change. To try it without writing code, you can also pick Kolibri-1 in the model selector of LLMTR Chat during the free period.
Frequently asked questions
Do I need to top up to use Kolibri-1?
No. An LLMTR account and an API key are enough. Requests are not charged to your balance; the model is offered free with a daily usage quota.
Why is Kolibri-1 limited to 32,768 tokens when the model card says 262,144?
262,144 tokens is the context length the model was trained with. In the free offering on LLMTR, input and output together must fit in 32,768 tokens; one response returns at most 16,384 tokens.
What happens when the free period ends?
After 9 October 2026, 23:59 (TRT), requests get a 410 model_retired response. No replacement is suggested because the catalog has no other Aleph Alpha model; the model page stays open.