Integration guides · 2026-10-04

Kolibri-1 tool calling and reasoning_effort: a developer guide for the LLMTR API

A step-by-step guide to Kolibri-1 on the LLMTR API: choosing the reasoning level, fitting the max_tokens budget into the 32,768-token window, defining tools and building the tool turn correctly.

LLMTR diagram showing the reasoning_effort levels, the max_tokens budget and the steps of a tool-calling turn for Kolibri-1.

Short answer

Reasoning is off by default on tesseracted/kolibri-1. When you set reasoning_effort to low, medium or high the model reasons before it answers; the tokens spent on reasoning count toward max_tokens. Tool calling works: the model picks a tool, returns the arguments as JSON and answers from the tool result.

Watch two things. The input plus max_tokens must fit in 32,768 tokens, and every tool call that sits in the history must have a tool result. The sections below give those two rules and example request bodies. The model id and the terms of the free period are covered in the announcement article.

reasoning_effort: none, low, medium, high

You can choose the level with the reasoning_effort field in the request body or with a suffix on the model id: tesseracted/kolibri-1:high. If both are sent, the field in the body wins. If you omit the field or send none, the model answers without reasoning.

minimal, xhigh and max do not exist on this model; a request that sends one of them, or an unknown suffix, gets a 400. The length of the reasoning varies from request to request even at the same level. When choosing between medium and high, compare them on your own task; for short classification and labelling work, leaving reasoning off is faster.

reasoning_effort levels for tesseracted/kolibri-1
LevelSuffixBehaviour
none (default):none or no suffixReasoning off, direct answer
low:lowReasoning on, lowest level
medium:mediumReasoning on, middle level
high:highReasoning on, highest level
minimal, xhigh, max-Not supported, 400 response

Reasoning counts toward the output limit: the max_tokens budget

Reasoning tokens are counted inside completion_tokens and are spent out of max_tokens. If you do not send max_tokens the limit is 4,096 tokens. With reasoning on, a small max_tokens cuts the response off during the reasoning stage: finish_reason is length and content comes back empty. When you get such a response, raise max_tokens first or lower the level.

One response returns at most 16,384 tokens, and the input plus max_tokens cannot exceed 32,768 tokens. If you work with a long history, size the room you leave for max_tokens accordingly. When the limit is exceeded the request gets a 400; the message says how many tokens the input uses. The reasoning text is returned in the message.reasoning_content field of the response.

POST /v1/chat/completions: reasoning on, with room left for the answer

{
  "model": "tesseracted/kolibri-1",
  "messages": [
    {
      "role": "user",
      "content": "Ein Zug fährt um 9:40 Uhr ab und braucht 2 Stunden 35 Minuten. Wann kommt er an? Begründe kurz."
    }
  ],
  "reasoning_effort": "high",
  "max_tokens": 4000
}

Rules for tool definitions

Only tools of type function are defined in the tools field. Tool names must be unique and consist of 1 to 64 letters, digits, underscores or hyphens. A request can carry at most 64 tools, and the definitions cannot exceed 30,000 characters in total as JSON.

The parameters field must be a JSON Schema whose type is object; wrap a tool that expects an array or plain text in an object with a single property. The strict field has no effect on this model: arguments are not guaranteed to match the schema, so validate them on your side before running the tool. If you force a specific tool with tool_choice, that tool has to be defined in tools.

Build the tool turn correctly

When the model calls a tool, the response carries tool_calls and finish_reason is tool_calls. In the next request, add that assistant message to the history as it is, then send one tool message per call whose tool_call_id matches the id of that call. If a user, system or assistant message comes in before every call is answered, the request gets a 400. Call ids must be unique across the conversation.

When you trim the history, remove a tool call and its result together; leaving only one of them makes the request invalid. You can send the reasoning_content field of the assistant message back; that text counts toward the 60,000-character history limit. For a turn that must not call a tool, leave out the tools field instead of sending tool_choice none; with tools defined, tool_choice none gets a 400.

POST /v1/chat/completions: second turn carrying the tool result

{
  "model": "tesseracted/kolibri-1",
  "messages": [
    {
      "role": "user",
      "content": "Wie ist das Wetter in Berlin?"
    },
    {
      "role": "assistant",
      "content": null,
      "tool_calls": [
        {
          "id": "call_1",
          "type": "function",
          "function": {
            "name": "get_weather",
            "arguments": "{\"city\":\"Berlin\"}"
          }
        }
      ]
    },
    {
      "role": "tool",
      "tool_call_id": "call_1",
      "content": "{\"temp_c\":19}"
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Returns the current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string"
            }
          },
          "required": [
            "city"
          ]
        }
      }
    }
  ],
  "max_tokens": 1000
}

Streaming and token counts

When you send stream true, the reasoning text arrives in delta.reasoning_content and the answer in delta.content; tool calls are also delivered in fragments. To see the token count while streaming, set stream_options.include_usage to true: usage arrives on the final chunk. Reasoning tokens are included in completion_tokens.

Fields the model does not know do not fail the request. Fields such as seed, frequency_penalty, presence_penalty or user are not passed to the model and are listed in the llmtr_dropped_parameters field of the response; when streaming, that field arrives on the final chunk. If you are writing a test that relies on a repeatable answer, take into account that seed is not applied on this model.

Free period and thanks

Kolibri-1 is free on the LLMTR API until 9 October 2026 at 23:59 (Türkiye time); after that, requests get a 410 model_retired response. The free period is made possible by the support of Tesseracted Labs, and we thank them for it.

Frequently asked questions

Is reasoning on by default on Kolibri-1?

No. If reasoning_effort is omitted or set to none, the model answers without reasoning. Choosing low, medium or high turns reasoning on, and its tokens count toward max_tokens.

Can I get tool arguments that match the schema exactly with strict: true?

No. The strict field has no effect on this model and arguments are not guaranteed to match the schema. Validate them on your side before running the tool.

Why do I get a 400 when I send tool_choice none with tools defined?

On this model tool_choice none is not supported while tools are defined, and the request gets a 400. Leave the tools field out of the request for a turn that must not call a tool; the model then answers with text only.

Related posts