GLM API pricing: 2 models, per token and per task
The list price of every Z.ai model on our list, read from Z.ai’s own pricing page, next to what a passing answer cost on the AIGROW small-business benchmark and what three everyday jobs cost on each.
US dollars per million tokens, at the standard rate. No reseller prices.
- GLM-5.3
- $1.40 / $4.40
- GLM-5.3-Flash
- $0.15 / $0.50
Input / output list price per million tokens.
Z.ai prices the 2 models on our list from GLM-5.3-Flash at $0.15 in and $0.50 out per million tokens to GLM-5.3 at $1.40 in and $4.40 out. On the AIGROW small-business benchmark GLM-5.3-Flash shared the best record, 18 of 20 tasks passed, at the lowest cost: $0.831 per thousand passes.
Per token and per pass
The list price per million tokens, then how many of the benchmark’s 20 office tasks each model passed and what 1,000 passing answers cost, failed runs included.
| Model | Input | Output | Cached input | Context | Tasks passed | 1,000 passes |
|---|---|---|---|---|---|---|
| GLM-5.3Flagships | $1.40 | $4.40 | $0.26 | 1M | 18 of 20 | $7.02 |
| GLM-5.3-FlashSmall and fast | $0.15 | $0.50 | $0.03 | 1M | 18 of 20 | $0.831 |
What three everyday jobs cost
1,000 runs of each job at list price. The token counts are ours; the prices are Z.ai’s.
| Model | 1,000 × support reply | 1,000 × document summary | 1,000 × agent task |
|---|---|---|---|
| GLM-5.3 | $5.96 | $39.40 | $45.40 |
| GLM-5.3-Flash | $0.65 | $4.25 | $5.10 |
- Support reply
- A 3,000-token prompt (the ticket, the thread so far and a policy excerpt) and a 400-token answer.
- Document summary
- A 40-page document of about 25,000 tokens in, a 1,000-token summary out.
- Agent task
- Ten steps that read 80,000 tokens between them, 70,000 of them context repeated from step to step and served from cache, and write 3,000.
The small print
What Z.ai’s pricing pages add to the headline rates.
- GLM-5.3
- Cached-input storage is 'Limited-time Free'. Max output 128K. Open weights on Hugging Face (zai-org/GLM-5.3).
- GLM-5.3-Flash
- 320B total, 18B active parameters. Cached-input storage is 'Limited-time Free'. Z.ai also sells GLM-5.3-FlashX ($0.37 / $1.25), not included. Open weights on Hugging Face (zai-org/GLM-5.3-Flash).
Head to head
The comparisons with a GLM model on one side, each with both prices and the three jobs costed on each.
Where the prices come from
Read on 22 September 2026. Each price is the standard rate for the shortest context band, with no batch or priority discount; long-context surcharges, batch rates and promotions are in each model’s small print. No price was taken from a reseller.
Pay a quarter of list price
AIGROW API credit is metered at the makers’ list prices, like the ones above, and $25 buys $100 of usage. The credit page lists the models it covers.
Prices checked 22 September 2026