GLM-5.3-Flash: what it costs to run
The price Z.ai lists for GLM-5.3-Flash, what a support reply, a document summary and an agent task cost at it, and where that puts it among 37 models.
US dollars per million tokens, standard rate, shortest context band.
- 1,000 support replies
- $0.65
- Cached input
- $0.03
- Context window
- 1,000,000 tokens
- Released
- 26 Aug 2026
- Weights
- Open
From Z.ai’s own pages, 22 September 2026.
GLM-5.3-Flash costs $0.15 per million input tokens and $0.50 per million output tokens on Z.ai’s own API, and $0.03 for cached input. At those prices 1,000 support replies cost $0.65, the 2nd cheapest of 33 priced models on our list.
What a job costs
Three jobs, 1,000 runs of each, at the list price above, and where each puts GLM-5.3-Flash among the 33 models with a list price.
| Job | 1,000 runs | Among 33 |
|---|---|---|
| Support reply3,000 in, 400 out. | $0.65 | 2nd cheapest |
| Document summary25,000 in, 1,000 out. | $4.25 | 2nd cheapest |
| Agent task80,000 in (70,000 cached), 3,000 out. | $5.10 | cheapest |
The token counts are ours, chosen to look like real work; the prices are Z.ai’s. All three jobs stay under 200,000 tokens a request, below every long-context surcharge on the list.
GLM-5.3-Flash on real small-business tasks
What the AIGROW benchmark measured when GLM-5.3-Flash was given 20 office tasks, each marked pass or fail by fixed checks.
- Success rate
- 90% of 20 runs passed (95% interval 70 to 97%), 4th of 36 models.
- Cost per 1,000 passes
- $0.831, failed runs included.
- Answer time
- 15 s median, 31 s at the 90th percentile.
- Strongest and weakest
- Appointment scheduling (100% passed) and product descriptions (0%).
The small print
What the base rate above leaves out, copied from the maker’s pages: surcharges, discounts, promotions and cache pricing.
320B total, 18B active parameters. Cached-input storage is 'Limited-time Free'. Z.ai also sells GLM-5.3-FlashX ($0.37 / $1.25), not included. Open weights on Hugging Face (zai-org/GLM-5.3-Flash).
GLM-5.3-Flash against the models it is weighed with
Each comparison puts both list prices side by side and costs the same three jobs on each.
- vs Qwen3.8 Flash
- Two new Chinese budget models priced within a few cents of each other.
Other Z.ai models
| Model | Input | Output | Context | 1,000 replies |
|---|---|---|---|---|
| GLM-5.3Z.ai · open weights | $1.40 | $4.40 | 1M | $5.96 |
Where the price comes from
Read on 22 September 2026. Each price is the standard rate for the shortest context band, with no batch or priority discount; long-context surcharges, batch rates and promotions are in each model’s small print. No price was taken from a reseller.
Pay a quarter of list price
AIGROW API credit is metered at the makers’ list prices, and $25 buys $100 of usage. The credit page lists the models it covers; ask about any other before buying.
Prices checked 22 September 2026