Qwen / qwen/qwen3.8-2.4t-a95b
Qwen3.8 2.4T A95B - access through LLMTR
Qwen3.8 2.4T A95B is the open-weight release of the Qwen3.8 generation: a sparse Mixture-of-Experts model that runs 95 billion of its 2.4 trillion parameters per token. It is built for coding, research and long-horizon agent workflows that have to carry a multi-step task through to the end. It offers a 262,144-token context window, native tool calling, JSON Schema structured output and prompt caching that discounts repeated prefixes. Step-by-step reasoning is always on and returned separately from the answer, so a short question can still spend reasoning tokens; it accepts text only, with no image, audio or video input.
Technical specifications
| Canonical ID | qwen/qwen3.8-2.4t-a95b |
|---|---|
| Provider | Qwen |
| Context window | 262,144 tokens |
| Operations | CHAT_COMPLETIONS |
| Modalities | text |
Pricing
An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.
| Operation | Metric | Unit | Price |
|---|---|---|---|
| CHAT_COMPLETIONS | INPUT_TEXT | PER_1M_TOKENS | $2.00 |
| CHAT_COMPLETIONS | CACHE_READ | PER_1M_TOKENS | $0.200000 |
| CHAT_COMPLETIONS | OUTPUT_TEXT | PER_1M_TOKENS | $6.00 |
Example usage
With existing OpenAI SDK flows, change only the base URL and model identifier.
curl https://llmtr.com/v1/chat/completions \
-H "Authorization: Bearer llmtr-your_key" \
-H "Content-Type: application/json" \
-d '{"model":"qwen/qwen3.8-2.4t-a95b","messages":[{"role":"user","content":"Hello"}]}'
Related models
- Qwen3.8-Max qwen/qwen3.8-max
- Qwen3.7-Max qwen/qwen3.7-max
- Qwen3.7-Plus qwen/qwen3.7-plus
- Qwen3.6-Flash qwen/qwen3.6-flash
- Qwen-Plus qwen/qwen-plus
- Qwen-Plus-2025-01-25 qwen/qwen-plus-2025-01-25
- Qwen-Plus-2025-04-28 qwen/qwen-plus-2025-04-28
- Qwen-Plus-2025-07-14 qwen/qwen-plus-2025-07-14