DeepSeek V4.1 Flash: what it costs to run
The price DeepSeek lists for DeepSeek V4.1 Flash, what a support reply, a document summary and an agent task cost at it, and where that puts it among 37 models.
US dollars per million tokens, standard rate, shortest context band.
- 1,000 support replies
- $1.38
- Cached input
- $0.006
- Context window
- 1,000,000 tokens
- Released
- 10 Sep 2026
- Weights
- Open
From DeepSeek’s own pages, 22 September 2026.
DeepSeek V4.1 Flash costs $0.30 per million input tokens and $1.20 per million output tokens on DeepSeek’s own API, and $0.006 for cached input. At those prices 1,000 support replies cost $1.38, the 8th cheapest of 33 priced models on our list.
What a job costs
Three jobs, 1,000 runs of each, at the list price above, and where each puts DeepSeek V4.1 Flash among the 33 models with a list price.
| Job | 1,000 runs | Among 33 |
|---|---|---|
| Support reply3,000 in, 400 out. | $1.38 | 8th cheapest |
| Document summary25,000 in, 1,000 out. | $8.70 | 8th cheapest |
| Agent task80,000 in (70,000 cached), 3,000 out. | $7.02 | 3rd cheapest |
The token counts are ours, chosen to look like real work; the prices are DeepSeek’s. All three jobs stay under 200,000 tokens a request, below every long-context surcharge on the list.
DeepSeek V4.1 Flash on real small-business tasks
What the AIGROW benchmark measured when DeepSeek V4.1 Flash was given 20 office tasks, each marked pass or fail by fixed checks.
- Success rate
- 90% of 20 runs passed (95% interval 70 to 97%), 4th of 36 models.
- Cost per 1,000 passes
- $1.05, failed runs included.
- Answer time
- 14 s median, 127 s at the 90th percentile.
- Strongest and weakest
- Appointment scheduling (100% passed) and review replies (0%).
The small print
What the base rate above leaves out, copied from the maker’s pages: surcharges, discounts, promotions and cache pricing.
Peak-hour rate, recorded as the list price. Off-peak is half: $0.15 / $0.60, cache hit $0.003. API name deepseek-flash. Max output 384K.
DeepSeek V4.1 Flash against the models it is weighed with
Each comparison puts both list prices side by side and costs the same three jobs on each.
- vs gpt-oss-120b
- An open-weight model with no first-party price against the cheapest first-party API from DeepSeek.
- vs GPT-5.6 Luna
- Two low-cost APIs for bulk work, with DeepSeek cheaper again off-peak.
- vs Gemini 3.5 Flash-Lite
- Budget models with 1M context, compared for document-heavy tasks.
- vs Qwen3.8 Flash
- Chinese budget APIs, with Qwen's list price about half of DeepSeek's peak rate.
- vs Grok 4.3
- xAI's cheapest model against DeepSeek's, for teams comparing low-cost options outside OpenAI and Google.
Other DeepSeek models
| Model | Input | Output | Context | 1,000 replies |
|---|---|---|---|---|
| DeepSeek V4 ProDeepSeek · open weights | $1.32 | $3.96 | 1M | $5.54 |
Where the price comes from
Read on 22 September 2026. Each price is the standard rate for the shortest context band, with no batch or priority discount; long-context surcharges, batch rates and promotions are in each model’s small print. No price was taken from a reseller.
Run DeepSeek V4.1 Flash for a quarter of its list price
AIGROW API credit covers DeepSeek V4.1 Flash and is metered at the list price above: $25 buys $100 of usage.
Prices checked 22 September 2026