InclusionAI / inclusionai/ling-3.1-flash

Ling 3.1 Flash - access through LLMTR

Ling 3.1 Flash is InclusionAI's hybrid reasoning model: a Mixture-of-Experts with 560B total parameters, about 25B of them active per token. It offers a 256K context and returns at most 32,768 tokens per response. Reasoning is on by default; send `reasoning_effort: "none"` (or the `:none` suffix) to turn it off - that is the only level this model acts on. Tool calling works in either reasoning state. Repeated prefixes are served from a prompt cache; since the card is free the gain is latency, not price. It accepts text only - no image or audio input, and no schema-enforced JSON output. The free period ends on 13 October 2026 at 19:00 Turkey time, after which requests return an error naming its successor, the metered `inclusionai/ling-3.0-flash`.

Technical specifications

Canonical IDinclusionai/ling-3.1-flash
ProviderInclusionAI
Context window262,144 tokens
OperationsCHAT_COMPLETIONS
Modalitiestext

Pricing

An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.

OperationMetricUnitPrice
CHAT_COMPLETIONSINPUT_TEXTPER_1M_TOKENSNot available
CHAT_COMPLETIONSCACHE_READPER_1M_TOKENSNot available
CHAT_COMPLETIONSOUTPUT_TEXTPER_1M_TOKENSNot available

Example usage

With existing OpenAI SDK flows, change only the base URL and model identifier.

curl https://llmtr.com/v1/chat/completions   -H "Authorization: Bearer llmtr-your_key"   -H "Content-Type: application/json"   -d '{"model":"inclusionai/ling-3.1-flash","messages":[{"role":"user","content":"Hello"}]}'

Related models