Qwen / qwen/qwen3.8-2.4t-a95b

Qwen3.8 2.4T A95B - access through LLMTR

Qwen3.8 2.4T A95B is the open-weight release of the Qwen3.8 generation: a sparse Mixture-of-Experts model that runs 95 billion of its 2.4 trillion parameters per token. It is built for coding, research and long-horizon agent workflows that have to carry a multi-step task through to the end. It offers a 262,144-token context window, native tool calling, JSON Schema structured output and prompt caching that discounts repeated prefixes. Step-by-step reasoning is always on and returned separately from the answer, so a short question can still spend reasoning tokens; it accepts text only, with no image, audio or video input.

Technical specifications

Canonical IDqwen/qwen3.8-2.4t-a95b
ProviderQwen
Context window262,144 tokens
OperationsCHAT_COMPLETIONS
Modalitiestext

Pricing

An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.

OperationMetricUnitPrice
CHAT_COMPLETIONSINPUT_TEXTPER_1M_TOKENS$2.00
CHAT_COMPLETIONSCACHE_READPER_1M_TOKENS$0.200000
CHAT_COMPLETIONSOUTPUT_TEXTPER_1M_TOKENS$6.00

Example usage

With existing OpenAI SDK flows, change only the base URL and model identifier.

curl https://llmtr.com/v1/chat/completions \
  -H "Authorization: Bearer llmtr-your_key" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen/qwen3.8-2.4t-a95b","messages":[{"role":"user","content":"Hello"}]}'

Related models