DeepSeek V4.1 Flash vs GPT-5.6 Luna
Two low-cost APIs for bulk work, with DeepSeek cheaper again off-peak. Both list prices from the makers’ own pages, and what three everyday jobs cost on each.
US dollars, standard rate, shortest context band. Price only: this page does not rank quality.
- DeepSeek V4.1 Flash
- $1.38
- GPT-5.6 Luna
- $1.08
- DeepSeek V4.1 Flash, per M tokens
- $0.30 / $1.20
- GPT-5.6 Luna, per M tokens
- $0.20 / $1.20
3,000 tokens in and 400 out each, at list price.
Per million tokens, DeepSeek V4.1 Flash costs $0.30 in and $1.20 out; GPT-5.6 Luna costs $0.20 in and $1.20 out. 1,000 support replies cost $1.38 on DeepSeek V4.1 Flash and $1.08 on GPT-5.6 Luna, so DeepSeek V4.1 Flash costs 28% more. DeepSeek V4.1 Flash reads up to 1M tokens at once, GPT-5.6 Luna up to 1.05M.
Side by side
Each maker’s published rate for its standard tier. Cached input is what a repeated prompt prefix costs once the provider has stored it.
| Fact | DeepSeek V4.1 Flash | GPT-5.6 Luna |
|---|---|---|
| Maker | DeepSeek | OpenAI |
| Tier | Small and fast | Small and fast |
| Input, per M tokens | $0.30 | $0.20 |
| Output, per M tokens | $1.20 | $1.20 |
| Cached input | $0.006 | $0.02 |
| Context window | 1M | 1.05M |
| Released | 10 Sep 2026 | 9 Jul 2026 |
| Weights | Open | Closed |
What a job costs on each
1,000 runs of each job at the two list prices. The gap moves from job to job because the models price input, output and cached context differently.
| Job | DeepSeek V4.1 Flash | GPT-5.6 Luna | Cheaper |
|---|---|---|---|
| Support reply3,000 in, 400 out. | $1.38 | $1.08 | GPT-5.6 Luna, 22% less |
| Document summary25,000 in, 1,000 out. | $8.70 | $6.20 | GPT-5.6 Luna, 29% less |
| Agent task80,000 in (70,000 cached), 3,000 out. | $7.02 | $7.00 | Even |
The token counts are ours, chosen to look like real work; the prices are the makers’. A model with no published cache price bills the agent task’s repeated context at its full input rate.
The small print
What each base rate leaves out, from the makers’ own pages: surcharges, discounts, promotions and cache pricing.
- DeepSeek V4.1 Flash
- Peak-hour rate, recorded as the list price. Off-peak is half: $0.15 / $0.60, cache hit $0.003. API name deepseek-flash. Max output 384K. DeepSeek’s pricing page
- GPT-5.6 Luna
- Nano tier of the 5.6 family. Price cut 80% on 2026-07-30. Requests above 272K input tokens bill 2x input and 1.5x output ($0.40 / $1.80). Cache write $0.25. Batch and Flex 50% off. Data residency adds 10%. OpenAI’s pricing page
Other comparisons
Every head to head on the list that involves DeepSeek V4.1 Flash or GPT-5.6 Luna.
- gpt-oss-120b vs DeepSeek V4.1 Flash
- GPT-5.6 Luna vs Gemini 3.5 Flash-Lite
- GPT-5.6 Luna vs Gemini 3.1 Flash-Lite
- Claude Haiku 4.5 vs GPT-5.6 Luna
- GPT-5.4 mini vs GPT-5.6 Luna
- GPT-5.4 nano vs GPT-5.6 Luna
- DeepSeek V4.1 Flash vs Gemini 3.5 Flash-Lite
- Qwen3.8 Flash vs DeepSeek V4.1 Flash
- Mistral Small 4 vs GPT-5.6 Luna
- Grok 4.3 vs DeepSeek V4.1 Flash
Run both for a quarter of list price
AIGROW API credit covers DeepSeek V4.1 Flash and GPT-5.6 Luna and is metered at the list prices above: $25 buys $100 of usage.
Prices checked 22 September 2026