Qwen / qwen/qwen3.8-flash

Qwen3.8-Flash - access through LLMTR

Qwen3.8-Flash is Qwen's low-cost Flash model for applications that examine long documents, images and videos in the same conversation. It provides a 1,000,000-token context and a 131,072-token output limit. It supports coding assistance, document summaries, visual question answering and tool-using workflows, with JSON Schema for structured text output. Thinking is on by default and can be disabled. Cache hits on repeated prompt prefixes reduce input cost; long inputs use the same pricing tier.

Technical specifications

Canonical IDqwen/qwen3.8-flash
ProviderQwen
Context window1,000,000 tokens
OperationsCHAT_COMPLETIONS
Modalitiestext, image, video

Pricing

An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.

OperationMetricUnitPrice
CHAT_COMPLETIONSINPUT_TEXTPER_1M_TOKENS$0.160000
CHAT_COMPLETIONSCACHE_READPER_1M_TOKENS$0.016000
CHAT_COMPLETIONSOUTPUT_TEXTPER_1M_TOKENS$0.470000

Example usage

With existing OpenAI SDK flows, change only the base URL and model identifier.

curl https://llmtr.com/v1/chat/completions \
  -H "Authorization: Bearer llmtr-your_key" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen/qwen3.8-flash","messages":[{"role":"user","content":"Hello"}]}'

Related models