GLM-5.3-Flash vs Qwen3.8 Flash
Two new Chinese budget models priced within a few cents of each other. Both list prices from the makers’ own pages, and what three everyday jobs cost on each.
US dollars, standard rate, shortest context band. Price only: this page does not rank quality.
- GLM-5.3-Flash
- $0.65
- Qwen3.8 Flash
- $0.64
- GLM-5.3-Flash, per M tokens
- $0.15 / $0.50
- Qwen3.8 Flash, per M tokens
- $0.15 / $0.47
3,000 tokens in and 400 out each, at list price.
Per million tokens, GLM-5.3-Flash costs $0.15 in and $0.50 out; Qwen3.8 Flash costs $0.15 in and $0.47 out. 1,000 support replies cost $0.65 on GLM-5.3-Flash and $0.64 on Qwen3.8 Flash, so GLM-5.3-Flash costs 2% more. Both read up to 1M tokens at once.
Side by side
Each maker’s published rate for its standard tier. Cached input is what a repeated prompt prefix costs once the provider has stored it.
| Fact | GLM-5.3-Flash | Qwen3.8 Flash |
|---|---|---|
| Maker | Z.ai | Alibaba (Qwen) |
| Tier | Small and fast | Small and fast |
| Input, per M tokens | $0.15 | $0.15 |
| Output, per M tokens | $0.50 | $0.47 |
| Cached input | $0.03 | Not published |
| Context window | 1M | 1M |
| Released | 26 Aug 2026 | Not published |
| Weights | Open | Closed |
What a job costs on each
1,000 runs of each job at the two list prices. The gap moves from job to job because the models price input, output and cached context differently.
| Job | GLM-5.3-Flash | Qwen3.8 Flash | Cheaper |
|---|---|---|---|
| Support reply3,000 in, 400 out. | $0.65 | $0.64 | Qwen3.8 Flash, 2% less |
| Document summary25,000 in, 1,000 out. | $4.25 | $4.22 | Even |
| Agent task80,000 in (70,000 cached), 3,000 out. | $5.10 | $13.41 | GLM-5.3-Flash, 62% less |
The token counts are ours, chosen to look like real work; the prices are the makers’. A model with no published cache price bills the agent task’s repeated context at its full input rate.
The small print
What each base rate leaves out, from the makers’ own pages: surcharges, discounts, promotions and cache pricing.
- GLM-5.3-Flash
- 320B total, 18B active parameters. Cached-input storage is 'Limited-time Free'. Z.ai also sells GLM-5.3-FlashX ($0.37 / $1.25), not included. Open weights on Hugging Face (zai-org/GLM-5.3-Flash). Z.ai’s pricing page
- Qwen3.8 Flash
- International deployment, one tier up to 1M input tokens. Cache discounts as for Qwen3.8 Max. Context taken from the top pricing tier (1M). Alibaba (Qwen)’s pricing page
Other comparisons
Every head to head on the list that involves GLM-5.3-Flash or Qwen3.8 Flash.
Pay a quarter of list price
AIGROW API credit is metered at the makers’ list prices, and $25 buys $100 of usage. The credit page lists the models it covers; ask about any other before buying.
Prices checked 22 September 2026