Z.AI / zai/glm-5.3-flashx
GLM-5.3-FlashX - access through LLMTR
GLM-5.3-FlashX is a 1M-context Z.AI model that processes images, text, and code together with function calling. It accepts the same inputs as GLM-5.3-Flash, generates tokens faster, and costs more per token; pick it for agent loops where latency matters more than unit price. Reasoning runs on every response and cannot be turned off - plain requests use low, and :high or :max go deeper.
Technical specifications
| Canonical ID | zai/glm-5.3-flashx |
|---|---|
| Provider | Z.AI |
| Context window | 1,000,000 tokens |
| Operations | CHAT_COMPLETIONS |
| Modalities | text, image |
Pricing
An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.
| Operation | Metric | Unit | Price |
|---|---|---|---|
| CHAT_COMPLETIONS | INPUT_TEXT | PER_1M_TOKENS | $0.333000 |
| CHAT_COMPLETIONS | CACHE_READ | PER_1M_TOKENS | $0.067500 |
| CHAT_COMPLETIONS | OUTPUT_TEXT | PER_1M_TOKENS | $1.125 |
Example usage
With existing OpenAI SDK flows, change only the base URL and model identifier.
curl https://llmtr.com/v1/chat/completions -H "Authorization: Bearer llmtr-your_key" -H "Content-Type: application/json" -d '{"model":"zai/glm-5.3-flashx","messages":[{"role":"user","content":"Hello"}]}'
Related models
- GLM-5.3-Flash zai/glm-5.3-flash
- GLM-5.3 zai/glm-5.3
- GLM-5.2 zai/glm-5.2
- GLM-5.1 zai/glm-5.1
- GLM-5 zai/glm-5
- GLM-5-Turbo zai/glm-5-turbo
- GLM-5V-Turbo zai/glm-5v-turbo
- GLM-4.7 zai/glm-4.7