Gemini 3.8 Flash: what it costs to run
The price Google lists for Gemini 3.8 Flash, what a support reply, a document summary and an agent task cost at it, and where that puts it among 37 models.
US dollars per million tokens, standard rate, shortest context band.
- 1,000 support replies
- $3.75
- Cached input
- $0.075
- Context window
- 1,048,576 tokens
- Released
- 2 Sep 2026
- Weights
- Closed
From Google’s own pages, 22 September 2026.
Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens on Google’s own API, and $0.075 for cached input. At those prices 1,000 support replies cost $3.75, the 13th cheapest of 33 priced models on our list.
What a job costs
Three jobs, 1,000 runs of each, at the list price above, and where each puts Gemini 3.8 Flash among the 33 models with a list price.
| Job | 1,000 runs | Among 33 |
|---|---|---|
| Support reply3,000 in, 400 out. | $3.75 | 13th cheapest |
| Document summary25,000 in, 1,000 out. | $22.50 | 13th cheapest |
| Agent task80,000 in (70,000 cached), 3,000 out. | $24.00 | 11th cheapest |
The token counts are ours, chosen to look like real work; the prices are Google’s. All three jobs stay under 200,000 tokens a request, below every long-context surcharge on the list.
Gemini 3.8 Flash on real small-business tasks
What the AIGROW benchmark measured when Gemini 3.8 Flash was given 20 office tasks, each marked pass or fail by fixed checks.
- Success rate
- 90% of 20 runs passed (95% interval 70 to 97%), 4th of 36 models.
- Cost per 1,000 passes
- $7.17, failed runs included.
- Answer time
- 11 s median, 16 s at the 90th percentile.
- Strongest and weakest
- Appointment scheduling (100% passed) and review replies (0%).
The small print
What the base rate above leaves out, copied from the maker’s pages: surcharges, discounts, promotions and cache pricing.
Launch price through 2026-12-31; from 2027-01-01 it becomes $1.50 / $7.50, cached $0.15. Batch and Flex 50% off; Priority $1.35 / $6.75. Cache storage $0.50 per MTok per hour (rising to $1.00).
Gemini 3.8 Flash against the models it is weighed with
Each comparison puts both list prices side by side and costs the same three jobs on each.
- vs Claude Sonnet 5
- Anthropic's workhorse against Google's, where Flash stays cheaper even after its 2027 price rise.
- vs GPT-5.6 Terra
- OpenAI's and Google's mid tiers, the usual shortlist for high-volume business tasks.
- vs Qwen3.7 Plus
- Two low-cost mid-tier models with long context, suited to bulk text work.
- vs Mistral Medium 3.5
- The European mid-priced model against Google's high-volume workhorse.
- vs Grok 4.3
- xAI's low-cost model against Google's mid tier, with cheaper output on the Grok side.
- vs Claude Haiku 4.5
- Fast models from Anthropic and Google weighed for customer-facing replies, with Flash cheaper on both input and output.
Other Google models
| Model | Input | Output | Context | 1,000 replies |
|---|---|---|---|---|
| Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1.05M | $1.35 |
| Gemini 3.5 Flash-LiteGoogle | $0.30 | $2.50 | 1.05M | $1.90 |
| Gemini 3.1 Pro (preview)Google | $2.00 | $12.00 | 1.05M | $10.80 |
| Gemma 4 31BGoogle · open weights | No list price | · | 256K | · |
Where the price comes from
Read on 22 September 2026. Each price is the standard rate for the shortest context band, with no batch or priority discount; long-context surcharges, batch rates and promotions are in each model’s small print. No price was taken from a reseller.
Run Gemini 3.8 Flash for a quarter of its list price
AIGROW API credit covers Gemini 3.8 Flash and is metered at the list price above: $25 buys $100 of usage.
Prices checked 22 September 2026