Integration guides · 2026-09-30

MiniMax M3.1 Flash Preview on the LLMTR API: free access, limits and first request

Through our partnership with MiniMax, the closed-beta MiniMax M3.1 Flash Preview is free on the LLMTR API until 6 October 2026, 23:59 (TRT). Access rule, first request, 1M context and reasoning_effort.

LLMTR diagram showing the model id, the request endpoint and the end of the free period for MiniMax M3.1 Flash Preview.

Short answer: free until 6 October through our MiniMax partnership

Through the partnership between LLMTR and MiniMax, MiniMax M3.1 Flash Preview is available on the LLMTR API from 30 September 2026 under the id minimax/minimax-m3.1-flash-preview. The model is free: requests are not charged to your balance. The free period ends on 6 October 2026 at 23:59 (Turkey time).

MiniMax's API documentation describes this release as a multimodal coding model with a 1,000,000-token context window and tunable thinking depth. According to the same page, MiniMax's public channels for it are, for now, only its own subscription plan and MiniMax Code. Because the model is in closed beta, its behaviour may change during this period.

There is one condition: to use the model, your account must have been topped up at least once. The smallest top-up is $5. The credit you add is not spent on this model; it stays in your account for other models.

Send your first request

Create an LLMTR API key and send the request to the /v1/chat/completions endpoint. If you use the OpenAI SDK, set the base URL to LLMTR and write the full id, including the provider prefix, in the model field.

The body below lowers the thinking depth to low; that is enough for a short coding task and the answer arrives faster. The usage field in the response shows prompt_tokens and completion_tokens; the tokens spent on thinking are counted inside completion_tokens.

First test body for POST /v1/chat/completions

{
  "model": "minimax/minimax-m3.1-flash-preview",
  "messages": [
    {
      "role": "system",
      "content": "Answer briefly and clearly."
    },
    {
      "role": "user",
      "content": "Find the bug in this Python function and write the fixed version: def average(x): return sum(x) / len(x) - 1"
    }
  ],
  "reasoning_effort": "low",
  "max_tokens": 8000
}

Limits and input types

The card offers a 1,000,000-token context window and produces at most 524,288 tokens per response. If you send a larger max_tokens value, the request is refused before it reaches the model and the error message states the limit.

Text and image input are accepted; send images as OpenAI-style image_url parts. Tool calling and JSON object mode work. Schema-conforming output (json_schema) is not guaranteed, so validate the response on your side. Video input appears in MiniMax's documentation but is not listed on this card; for video, see minimax/minimax-m3.

Limits of the minimax/minimax-m3.1-flash-preview card
FeatureValue
Context window1,000,000 tokens
Maximum output per response524,288 tokens
InputText and images
Tool calling and JSON object modeSupported
reasoning_effortlow, medium, high, xhigh, max
Access ruleAccount topped up at least once (minimum $5)
End of free period6 October 2026, 23:59 (TRT)

Thinking is always on: reasoning_effort and max_tokens

The model thinks before every answer and this step cannot be turned off. With the reasoning_effort field you choose the thinking depth from low, medium, high, xhigh and max. According to MiniMax's documentation, if you omit the field the model runs at max, the deepest level; higher levels mean more thinking tokens and longer waits. Sending none gets the request refused with a 400.

Thinking tokens count toward the output limit. If you set max_tokens very low, the answer is cut off during thinking and finish_reason shows length. Leave a few thousand tokens even for short answers; if the answer comes back empty, first raise max_tokens or choose a lower reasoning_effort.

Access rule and error responses

If your account has never been topped up, the request gets a 402 response and the error.details.reason field in the error body says top_up_required. After you top up, wait a minute and send the request again. Credit granted at sign-up, through a campaign or as a gift does not meet the condition; accounts on an enterprise contract can use the model with no extra step.

Free models on LLMTR come with a daily usage quota. When it runs out you get a 429 and can try again the next day. If you send many requests in a short time you may also see a temporary 429; in that case use exponential backoff in your client. Because the model is in closed beta, it may also answer 503 model_unavailable for a while before the end date. On free models prompts may be logged by the provider, so do not send personal or confidential data.

What happens after 6 October

From 7 October 2026, requests to minimax/minimax-m3.1-flash-preview get a 410 model_retired response, and the response names the metered minimax/minimax-m3. The switch is not automatic: no request reaches the paid model and nothing is charged until you change the model id yourself. Your application should not retry a 410 as if it were a transient error.

minimax/minimax-m3 is the earlier release of the same family, and its price is listed on the pricing page. JSON mode and reasoning_effort are not listed on that card, so check requests that use response_format or reasoning_effort before you switch. To try the model without writing code, you can also pick it from the model selector in LLMTR Chat; the same access rule applies there.

Frequently asked questions

Does MiniMax M3.1 Flash Preview cost anything?

No. Requests are not charged to your balance. To use it, your account only needs to have been topped up once; the smallest top-up is $5 and the credit is used for other models.

I get a 402 top_up_required response. What should I do?

Add at least $5 to your account and repeat the request after a minute. Sign-up, campaign or gift credit does not meet the condition.

What happens when the free period ends?

After 6 October 2026, 23:59 (TRT), requests get a 410 model_retired response and minimax/minimax-m3 is suggested.

Related posts