Integration guides · 2026-09-28
How to use MiniMax M3.1 for free: first API request and limits
MiniMax M3.1 is free on LLMTR until 30 September 2026, 23:59 (TRT). Model id, first API request, 1M context, 32,768 output cap, the reasoning budget and quota behaviour.
Short answer: free until 30 September 2026, 23:59
On LLMTR, MiniMax M3.1 is offered free under the id minimax/minimax-m3.1-free. The free period ends on 30 September 2026 at 23:59 Turkey time (UTC+3). From 1 October, requests to this id receive a 410 model_retired response that names the metered minimax/minimax-m3. Nothing switches over on its own, and nothing is charged until you change the model id yourself.
MiniMax's own API documentation describes this release as MiniMax-M3.1-Flash-Preview: a coding model with a 1,000,000-token context window and tunable thinking depth. According to the same page, MiniMax currently offers it only through its Token Plan and MiniMax Code. The LLMTR card lets you try it with a standard API key over an OpenAI-compatible endpoint.
Send the first request
Create an LLMTR API key and send the request to the /v1/chat/completions endpoint. With the OpenAI SDK, set the base URL to LLMTR and put the full id in the model field. Write the id in full, with its provider prefix: minimax/minimax-m3.1-free.
The body below is enough for a small code review task. When the response arrives, look at prompt_tokens and completion_tokens in the usage field. The amount is zero on this free row, but the token counts are the most reliable way to estimate cost if you move to a metered model later.
First request body for POST /v1/chat/completions
{
"model": "minimax/minimax-m3.1-free",
"messages": [
{
"role": "system",
"content": "Answer briefly and precisely."
},
{
"role": "user",
"content": "Find the bug in this Python function and write the fixed version: def average(x): return sum(x) / len(x) - 1"
}
],
"max_tokens": 4000
}
Context, output and input types
The card offers a 1,000,000-token context window, and a single response returns at most 32,768 tokens. A larger max_tokens value is refused before it reaches the model, and the error message states the limit. If you need a longer output, split the work into parts and request each part separately.
Text and image input are accepted. Send an image as an OpenAI-style image_url part. Tool calling, JSON object mode and schema-enforced output work. Video input is not offered on this card; if you need video, look at the metered minimax/minimax-m3 row.
| Property | Value |
|---|---|
| Context window | 1,000,000 tokens |
| Maximum output per response | 32,768 tokens |
| Input | Text and images |
| Tool calling and JSON | Supported |
| End of the free period | 30 September 2026, 23:59 (TRT) |
Thinking is always on: do not set max_tokens low
M3.1 thinks before it answers, and that step cannot be switched off. MiniMax's documentation also states that this release always thinks and rejects a request to disable thinking. Thinking tokens count toward the output limit. If you set max_tokens very low, such as 20 or 50, the response is cut off during thinking and finish_reason reads length.
Leave a few hundred tokens even for short answers, and a few thousand for code and analysis tasks. If the answer comes back empty with finish_reason length, raise this value first. Repeating the same request unchanged does not change the outcome.
Quota and error responses
Free models on LLMTR come with a daily usage quota, and this card follows the same rule. When the quota is used up you receive a 429 response and can try again the next day. You may also see a temporary 429 when you send many requests in a short time. Wait a minute, retry, and use exponential backoff in your client.
Free access is limited, and the card may answer 503 model_unavailable before its end date. In that case, try again later. Do not put production traffic on this card; choose a metered model for work that must run without interruption. On free models the provider may log prompts, so do not send personal or confidential data.
What to do after 30 September
Keep the model id in configuration so that switching is usually a one-line change. If you use response_format, take care: JSON mode and schema-enforced output are not listed on the M3 card. After 1 October, if a 410 model_retired response arrives, your application should not retry it like a transient network error. Show the user that the free period has ended and choose the next model deliberately.
To stay in the same family, minimax/minimax-m3 is available as a metered model. It is the release before M3.1, and its price is listed on the pricing page. Before switching, run the same small task on both models and compare the results. To try the model without writing code, pick MiniMax M3.1 from the model selector in LLMTR Chat.
Frequently asked questions
Until when is MiniMax M3.1 free on LLMTR?
Until 30 September 2026 at 23:59 Turkey time. From 1 October the minimax/minimax-m3.1-free id answers 410 model_retired.
The answer is empty and finish_reason is length. Why?
The model thinks before every answer and those tokens count toward the output limit. Raise max_tokens; leave a few hundred tokens even for short answers.
Am I switched to the paid model when the free period ends?
No. Requests receive a 410 response naming minimax/minimax-m3, and nothing is charged until you change the model id yourself.