Qwen / qwen/qwen3.8-max

Qwen3.8-Max - access through LLMTR

Qwen3.8-Max understands image and video input alongside text. It is a fit for long-horizon agent tasks, multi-step coding work, analysis of long documents, and tool-using workflows. It offers a 1M context window and produces at most 128K tokens per response. Thinking is enabled by default, while latency-sensitive requests can turn it off with the model suffix or the request body.

Technical specifications

Canonical IDqwen/qwen3.8-max
ProviderQwen
Context window1,000,000 tokens
OperationsCHAT_COMPLETIONS
Modalitiestext, image, video

Pricing

An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.

OperationMetricUnitPrice
CHAT_COMPLETIONSINPUT_TEXTPER_1M_TOKENS$2.00
CHAT_COMPLETIONSCACHE_READPER_1M_TOKENS$0.250000
CHAT_COMPLETIONSOUTPUT_TEXTPER_1M_TOKENS$6.00

Example usage

With existing OpenAI SDK flows, change only the base URL and model identifier.

curl https://llmtr.com/v1/chat/completions \
  -H "Authorization: Bearer llmtr-your_key" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen/qwen3.8-max","messages":[{"role":"user","content":"Hello"}]}'

Related models