InclusionAI / inclusionai/ling-3.1-flash
Ling 3.1 Flash - access through LLMTR
Ling 3.1 Flash is InclusionAI's hybrid reasoning model: a Mixture-of-Experts with 560B total parameters, about 25B of them active per token. It offers a 256K context and returns at most 32,768 tokens per response. Reasoning is on by default; send `reasoning_effort: "none"` (or the `:none` suffix) to turn it off - that is the only level this model acts on. Tool calling works in either reasoning state. Repeated prefixes are served from a prompt cache; since the card is free the gain is latency, not price. It accepts text only - no image or audio input, and no schema-enforced JSON output. The free period ends on 13 October 2026 at 19:00 Turkey time, after which requests return an error naming its successor, the metered `inclusionai/ling-3.0-flash`.
Technical specifications
| Canonical ID | inclusionai/ling-3.1-flash |
|---|---|
| Provider | InclusionAI |
| Context window | 262,144 tokens |
| Operations | CHAT_COMPLETIONS |
| Modalities | text |
Pricing
An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.
| Operation | Metric | Unit | Price |
|---|---|---|---|
| CHAT_COMPLETIONS | INPUT_TEXT | PER_1M_TOKENS | Not available |
| CHAT_COMPLETIONS | CACHE_READ | PER_1M_TOKENS | Not available |
| CHAT_COMPLETIONS | OUTPUT_TEXT | PER_1M_TOKENS | Not available |
Example usage
With existing OpenAI SDK flows, change only the base URL and model identifier.
curl https://llmtr.com/v1/chat/completions -H "Authorization: Bearer llmtr-your_key" -H "Content-Type: application/json" -d '{"model":"inclusionai/ling-3.1-flash","messages":[{"role":"user","content":"Hello"}]}'
Related models
- Ling 3.0 Flash inclusionai/ling-3.0-flash
- Gemma 4 llmtr/gemma-4
- Qwen 3.6 35B-A3B llmtr/qwen3-6-35b
- Trendyol Asure 12B llmtr/trendyol-asure-12b
- Magibu 11B v8 llmtr/magibu-11b-v8
- Muse Glimmer 30B (Turkey) llmtr/muse-glimmer-30b-tr
- Leyla llmtr/leyla
- EmbeddingGemma 300M llmtr/embeddinggemma-300m