InclusionAI / inclusionai/ling-3.0-flash
Ling 3.0 Flash - access through LLMTR
Ling 3.0 Flash is InclusionAI's 124-billion-parameter Mixture-of-Experts model, activating roughly 5.1 billion parameters per token. It is tuned for token efficiency in agentic workloads: it supports native tool calling, reads repeated prefixes from a prompt cache, and offers a 256K context window. It reasons by default, returning that chain separately in `reasoning_content`, whose tokens are counted inside `completion_tokens`. To turn reasoning off, send `reasoning_effort: "none"` or append the `:none` suffix to the model id - that is the only level with a measured effect on this model, so the others are not offered even though the upstream accepts them. A single response can be at most 32,768 tokens. It accepts text only - no image or audio input, and no schema-enforced JSON output. With reasoning on, the model can think at length on some prompts and exhaust the budget; when you need a short, direct answer, `none` is both faster and cheaper.
Technical specifications
| Canonical ID | inclusionai/ling-3.0-flash |
|---|---|
| Provider | InclusionAI |
| Context window | 262,144 tokens |
| Operations | CHAT_COMPLETIONS |
| Modalities | text |
Pricing
An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.
| Operation | Metric | Unit | Price |
|---|---|---|---|
| CHAT_COMPLETIONS | INPUT_TEXT | PER_1M_TOKENS | $0.060000 |
| CHAT_COMPLETIONS | CACHE_READ | PER_1M_TOKENS | $0.012000 |
| CHAT_COMPLETIONS | OUTPUT_TEXT | PER_1M_TOKENS | $0.180000 |
Example usage
With existing OpenAI SDK flows, change only the base URL and model identifier.
curl https://llmtr.com/v1/chat/completions -H "Authorization: Bearer llmtr-your_key" -H "Content-Type: application/json" -d '{"model":"inclusionai/ling-3.0-flash","messages":[{"role":"user","content":"Hello"}]}'
Guides about this model
- How to calculate the cost of running Ling 3.0 Tiny locally - Downloadable weights do not mean zero operating cost. Use your own measurements to calculate monthly totals and cost per accepted task instead of assuming free compute.
- Running Ling 3.0 Flash locally: the Ollama and MLX story differs from Tiny - The Ling Tiny on Mac article covered the small model's experimental local path. Flash is from the same family but a much larger model; this article covers how that size difference changes the hardware picture.
- Tool calling and reasoning_effort mechanics on Ling 3.0 Flash - The Flash migration guide covers the general framework of the move. This article covers two of Flash's technical behaviors in depth: tool-calling modes and reasoning_effort mechanics.
- Putting Ling 3.0 Flash into production with vLLM: hardware and batch settings - The Ling Tiny vLLM article covers the model path and single-server setup. This article covers the hardware sharing and batch settings that come into play when moving that same path to production scale for Flash.
- Financial analysis with reasoning and forced tool calls on Ling 3.0 Flash Fin - Ling 3.0 Flash Fin was a model tuned for multi-step investment research; its free offer ended on 30 September 2026 and the row was retired. This article covers when to keep thinking on, when to turn it off, and when to force a tool call; the same flow works on its metered successor, Ling 3.0 Flash.
- Lowering repeated system-instruction cost on Ling 3.0 Flash with prompt cache - The three existing articles covered this model's local trial, tool-calling/reasoning_effort mechanics, and vLLM production deployment — all on the self-hosting axis. This article covers how prompt cache lowers cost when calling it through the LLMTR API instead.
- Ling Tiny alternatives: free Flash Fin versus paid Flash - Flash Fin was a free alternative after Ling Tiny; its free offer ended on 30 September 2026 and the row was retired. Paid Flash remains the successor named in LLMTR's retirement policy for both Tiny and Fin.
- Ling 3.0 Tiny: migrating the retired LLMTR API route to Flash - For anyone searching Ling 3.0 Tiny: what the model was, why the provider closed the route, why LLMTR does not silently reroute the request, and the two things that change when you migrate to Ling 3.0 Flash.
Related models
- Ling 3.1 Flash inclusionai/ling-3.1-flash
- Gemma 4 llmtr/gemma-4
- Qwen 3.6 35B-A3B llmtr/qwen3-6-35b
- Trendyol Asure 12B llmtr/trendyol-asure-12b
- Magibu 11B v8 llmtr/magibu-11b-v8
- Muse Glimmer 30B (Turkey) llmtr/muse-glimmer-30b-tr
- Leyla llmtr/leyla
- EmbeddingGemma 300M llmtr/embeddinggemma-300m