Integration guides · 2026-10-01
Using MiniMax M3.1 Flash Preview with Claude Code and coding agents
Through our partnership with MiniMax, MiniMax M3.1 Flash Preview is free until 6 October 2026. How to connect it to Claude Code and OpenAI-compatible coding agents through LLMTR: environment variables, the :low suffix, quota and switching.
Short answer: a base URL and the full model id
LLMTR is an AI API gateway that serves the OpenAI and Anthropic formats with one API key (llmtr.com). Through our partnership with MiniMax, the closed-beta MiniMax M3.1 Flash Preview is free here until 6 October 2026, 23:59 (TRT), under the id minimax/minimax-m3.1-flash-preview. Connecting a coding agent takes two things: the base URL for the format the agent speaks, and the full model id including the provider prefix. Claude Code speaks the Anthropic Messages format; the base URL is https://llmtr.com and requests go to /v1/messages. For agents that speak the OpenAI format, the base URL is https://llmtr.com/v1 and requests go to /v1/chat/completions. Both paths use the same API key, the same access rule and the same daily quota. This model is not served on /v1/responses; a tool that requires the Responses API gets a 400 unsupported_operation response.
| Agent type | Base URL | Endpoint | Model field |
|---|---|---|---|
| Claude Code (Anthropic format) | https://llmtr.com | /v1/messages | minimax/minimax-m3.1-flash-preview:low |
| OpenAI-compatible agents | https://llmtr.com/v1 | /v1/chat/completions | minimax/minimax-m3.1-flash-preview |
| Tools that require the Responses API | Not supported | /v1/responses | 400 unsupported_operation |
Setting up Claude Code
Set the environment variables below before you start Claude Code. ANTHROPIC_MODEL sets the main model and ANTHROPIC_DEFAULT_HAIKU_MODEL sets the small model Claude Code calls for background work; giving both the same id sends every request in the session to this model. In Windows PowerShell, set the same values as $env:ANTHROPIC_BASE_URL = "https://llmtr.com".
Claude Code does not know this id from its own model list, so it assumes a 200K context window and prints a warning; the CLAUDE_CODE_MAX_CONTEXT_TOKENS=1000000 line reports the model's real 1,000,000-token window. The model picker only lists ids that contain claude or anthropic, so do not look for this model in the /model menu; set it with ANTHROPIC_MODEL.
Environment variables that point Claude Code at LLMTR and MiniMax M3.1 Flash Preview
export ANTHROPIC_BASE_URL=https://llmtr.com
export ANTHROPIC_AUTH_TOKEN=llmtr-your_key
export ANTHROPIC_MODEL=minimax/minimax-m3.1-flash-preview:low
export ANTHROPIC_DEFAULT_HAIKU_MODEL=minimax/minimax-m3.1-flash-preview:low
export CLAUDE_CODE_MAX_CONTEXT_TOKENS=1000000
claude
Check the setup: the task we ran
On 1 October 2026 we started Claude Code 2.1.286 with these settings in an empty folder and asked it to write a Python file and run it. Claude Code created hello.py, ran it with python and reported the output correctly. Across two runs the task took 4 and 5 model turns and finished in 16 to 19 seconds; tool calls and tool results travelled over /v1/messages without problems.
In this test Claude Code's system prompt and tool definitions came to about 25,000 tokens and were sent again with every request. The model is free, so these tokens are not charged to your balance, but even a short task takes several requests and each request counts toward the daily free quota. The dollar estimate Claude Code shows at the end of a session comes from its own price table and does not apply to this model; you see real usage on the usage page of your LLMTR dashboard.
Thinking depth: the :low suffix
The model thinks before every answer. According to MiniMax's API documentation, when reasoning_effort is not sent the deepest level, max, is used. Claude Code's own thinking and effort settings are not passed to this model; the way to choose a level is a suffix at the end of the model id: :low, :medium, :high, :xhigh or :max. A suffix outside the list, such as :turbo, is refused with 400 invalid_request and the error message lists the valid levels.
Start with :low for short steps such as editing files and running commands. Try :high or :max for multi-step debugging or design decisions; higher levels mean more thinking tokens and longer waits. The thinking text is not returned as a thinking block in the Anthropic-format response; the response contains only text and tool_use blocks.
OpenAI-compatible coding agents
In an agent that offers an OpenAI-compatible provider setting, fill in three fields: the base URL https://llmtr.com/v1, your LLMTR API key and the model id minimax/minimax-m3.1-flash-preview. Write the id in full, including the minimax/ prefix.
If the agent lets you add fields to the request body, use the reasoning_effort field; if it does not, use the same suffix and write the id as minimax/minimax-m3.1-flash-preview:low. If both are sent, the reasoning_effort field wins. If the agent has a maximum output setting, leave a few thousand tokens: thinking tokens count toward that limit, and if it is too low the answer is cut off while the model is still thinking.
Access rule, quota and after 6 October
If your account has never been topped up, the agent's first request gets a 402 response; in the OpenAI-format error body the error.details.reason field is top_up_required. Add at least $5 and try again after a minute; the model is not charged to your balance. When the daily quota runs out you get a 429. On free models prompts may be logged by the provider, so consider this before opening a confidential codebase.
From 7 October 2026 this id returns 410 model_retired and the agent stops. To move to the metered minimax/minimax-m3, change the ANTHROPIC_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL values and remove the :low suffix, because the M3 card lists no reasoning_effort levels. The switch is not automatic: no request reaches the paid model until you change it.
Connecting MiniMax M3.1 Flash Preview to Claude Code
Steps for running Claude Code with MiniMax M3.1 Flash Preview through LLMTR's Anthropic-compatible /v1/messages endpoint.
- Create an API key. Create an API key in the LLMTR dashboard.
- Meet the access rule. Top up your account at least once; the smallest top-up is $5 and this model is not charged to your balance.
- Set the environment variables. Set ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL, ANTHROPIC_DEFAULT_HAIKU_MODEL and CLAUDE_CODE_MAX_CONTEXT_TOKENS.
- Check with a short task. Start Claude Code and ask it to write a file and run it.
Frequently asked questions
Does MiniMax M3.1 Flash Preview work with Claude Code?
Yes. It works with ANTHROPIC_BASE_URL=https://llmtr.com and ANTHROPIC_MODEL=minimax/minimax-m3.1-flash-preview:low. On 1 October 2026 Claude Code 2.1.286 completed a write-and-run task with these settings.
How do I change the thinking depth in Claude Code?
Change the suffix at the end of the model id: :low, :medium, :high, :xhigh or :max. Without a suffix MiniMax's default, max, is used.
Will I pay the cost Claude Code shows?
No. That figure is Claude Code's own estimate. This model is free until 6 October 2026, 23:59 (TRT).