Z.AI / zai/glm-5.3-flashx

GLM-5.3-FlashX - access through LLMTR

GLM-5.3-FlashX is a 1M-context Z.AI model that processes images, text, and code together with function calling. It accepts the same inputs as GLM-5.3-Flash, generates tokens faster, and costs more per token; pick it for agent loops where latency matters more than unit price. Reasoning runs on every response and cannot be turned off - plain requests use low, and :high or :max go deeper.

Technical specifications

Canonical IDzai/glm-5.3-flashx
ProviderZ.AI
Context window1,000,000 tokens
OperationsCHAT_COMPLETIONS
Modalitiestext, image

Pricing

An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.

OperationMetricUnitPrice
CHAT_COMPLETIONSINPUT_TEXTPER_1M_TOKENS$0.333000
CHAT_COMPLETIONSCACHE_READPER_1M_TOKENS$0.067500
CHAT_COMPLETIONSOUTPUT_TEXTPER_1M_TOKENS$1.125

Example usage

With existing OpenAI SDK flows, change only the base URL and model identifier.

curl https://llmtr.com/v1/chat/completions   -H "Authorization: Bearer llmtr-your_key"   -H "Content-Type: application/json"   -d '{"model":"zai/glm-5.3-flashx","messages":[{"role":"user","content":"Hello"}]}'

Related models