DeepSeek / deepseek/deepseek-v4-flash-vision-exp

DeepSeek V4 Flash Vision Exp - access through LLMTR

Experimental DeepSeek V4 Flash Vision model with a 1M-token context window and up to 384K maximum output. Supports tool calling and prompt caching. Image inputs are accepted both as base64 data URLs and as remote https links, and are tokenized by image dimensions into the input token count.

Technical specifications

Canonical IDdeepseek/deepseek-v4-flash-vision-exp
ProviderDeepSeek
Context window1,000,000 tokens
OperationsCHAT_COMPLETIONS
Modalitiestext, image

Pricing

An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.

OperationMetricUnitPrice
CHAT_COMPLETIONSINPUT_TEXTPER_1M_TOKENS$0.220000
CHAT_COMPLETIONSCACHE_READPER_1M_TOKENS$0.007000
CHAT_COMPLETIONSOUTPUT_TEXTPER_1M_TOKENS$0.660000

Example usage

With existing OpenAI SDK flows, change only the base URL and model identifier.

curl https://llmtr.com/v1/chat/completions \
  -H "Authorization: Bearer llmtr-your_key" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4-flash-vision-exp","messages":[{"role":"user","content":"Hello"}]}'

Related models