DeepSeek-V4 Flash
MIT Open SourceDeveloped by DeepSeek · Released 2025-01-01
What are the token costs and operational benchmarks for DeepSeek-V4 Flash?
Architectural Overview
Frontier foundation model developed by DeepSeek featuring MLA v3 Ultra-Sparse MoE architecture.
Optimal Production Use Cases
Ultra-low-cost open-weights reasoning and coding with 1M context at unprecedented token generation speeds.
Interactive Monthly Token Economics & ROI Forecaster
Model your expected production workload across prompt (input) and completion (output) tokens.
Estimated Monthly Spend
$0.14/1M in · $0.56/1M out
$0.18/1M in · $0.72/1M out
Head-to-Head Comparisons Involving DeepSeek-V4 Flash
DeepSeek-V4 Flash vs Claude 3.5 Haiku
DeepSeek-V4 Flash is 82.5% cheaper for input tokens ($0.14 vs. $0.80 per 1M tokens) and $0.56 vs. $4.00 for output tokens (5.7x cost difference). In terms of operational performance, Claude 3.5 Haiku delivers faster response latency with 140ms TTFT (10ms faster than DeepSeek-V4 Flash). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 40.6% for Claude 3.5 Haiku.
DeepSeek-V4 Flash vs Claude 3.5 Sonnet
DeepSeek-V4 Flash is 95.3% cheaper for input tokens ($0.14 vs. $3.00 per 1M tokens) and $0.56 vs. $15.00 for output tokens (21.4x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (170ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.
DeepSeek-V4 Flash vs Claude 3.7 Sonnet
DeepSeek-V4 Flash is 95.3% cheaper for input tokens ($0.14 vs. $3.00 per 1M tokens) and $0.56 vs. $15.00 for output tokens (21.4x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (500ms faster than Claude 3.7 Sonnet). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.
DeepSeek-V4 Flash vs Claude Opus 5
DeepSeek-V4 Flash is 97.2% cheaper for input tokens ($0.14 vs. $5.00 per 1M tokens) and $0.56 vs. $25.00 for output tokens (35.7x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (190ms faster than Claude Opus 5). Claude Opus 5 leads coding benchmarks at 82.4% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.
DeepSeek-V4 Flash vs Codestral 25.01
DeepSeek-V4 Flash is 53.3% cheaper for input tokens ($0.14 vs. $0.30 per 1M tokens) and $0.56 vs. $0.90 for output tokens (2.1x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (0ms faster than Codestral 25.01). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 44.2% for Codestral 25.01.
DeepSeek-V4 Flash vs Composer 2.5
DeepSeek-V4 Flash is 92.2% cheaper for input tokens ($0.14 vs. $1.80 per 1M tokens) and $0.56 vs. $7.20 for output tokens (12.9x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (30ms faster than Composer 2.5). Composer 2.5 leads coding benchmarks at 74.6% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.
DeepSeek-V4 Flash vs DeepSeek-R1
DeepSeek-V4 Flash is 74.5% cheaper for input tokens ($0.14 vs. $0.55 per 1M tokens) and $0.56 vs. $2.19 for output tokens (3.9x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (1650ms faster than DeepSeek-R1). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 49.2% for DeepSeek-R1.
DeepSeek-V4 Flash vs DeepSeek-V3
Both models share identical input pricing at $0.14 per 1M tokens. In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (190ms faster than DeepSeek-V3). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 42.0% for DeepSeek-V3.
DeepSeek-V4 Flash vs Fable 5
DeepSeek-V4 Flash is 93.0% cheaper for input tokens ($0.14 vs. $2.00 per 1M tokens) and $0.56 vs. $8.00 for output tokens (14.3x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (90ms faster than Fable 5). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 58.0% for Fable 5.
DeepSeek-V4 Flash vs Gemini 2.0 Flash
Gemini 2.0 Flash is 28.6% cheaper for input tokens ($0.10 vs. $0.14 per 1M tokens) and $0.40 vs. $0.56 for output tokens (1.4x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (230ms faster than Gemini 2.0 Flash). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 48.0% for Gemini 2.0 Flash.
DeepSeek-V4 Flash vs Gemini 3.7 Flash
Gemini 3.7 Flash is 42.9% cheaper for input tokens ($0.08 vs. $0.14 per 1M tokens) and $0.32 vs. $0.56 for output tokens (1.8x cost difference). In terms of operational performance, Gemini 3.7 Flash delivers faster response latency with 75ms TTFT (75ms faster than DeepSeek-V4 Flash). Gemini 3.7 Flash leads coding benchmarks at 68.2% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.
DeepSeek-V4 Flash vs GLM 5.3 Flash
DeepSeek-V4 Flash is 6.7% cheaper for input tokens ($0.14 vs. $0.15 per 1M tokens) and $0.56 vs. $0.50 for output tokens (1.1x cost difference). In terms of operational performance, GLM 5.3 Flash delivers faster response latency with 130ms TTFT (20ms faster than DeepSeek-V4 Flash). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 54.2% for GLM 5.3 Flash.
DeepSeek-V4 Flash vs OpenAI GPT-4o
DeepSeek-V4 Flash is 94.4% cheaper for input tokens ($0.14 vs. $2.50 per 1M tokens) and $0.56 vs. $10.00 for output tokens (17.9x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (130ms faster than OpenAI GPT-4o). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 38.8% for OpenAI GPT-4o.
DeepSeek-V4 Flash vs GPT-5.6 Luna
DeepSeek-V4 Flash is 22.2% cheaper for input tokens ($0.14 vs. $0.18 per 1M tokens) and $0.56 vs. $0.72 for output tokens (1.3x cost difference). In terms of operational performance, GPT-5.6 Luna delivers faster response latency with 90ms TTFT (60ms faster than DeepSeek-V4 Flash). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 48.5% for GPT-5.6 Luna.
DeepSeek-V4 Flash vs GPT-5.6 Sol
DeepSeek-V4 Flash is 98.2% cheaper for input tokens ($0.14 vs. $8.00 per 1M tokens) and $0.56 vs. $32.00 for output tokens (57.1x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (270ms faster than GPT-5.6 Sol). GPT-5.6 Sol leads coding benchmarks at 79.5% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.
DeepSeek-V4 Flash vs GPT-5.6 Terra
DeepSeek-V4 Flash is 90.7% cheaper for input tokens ($0.14 vs. $1.50 per 1M tokens) and $0.56 vs. $6.00 for output tokens (10.7x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (60ms faster than GPT-5.6 Terra). GPT-5.6 Terra leads coding benchmarks at 65.4% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.
DeepSeek-V4 Flash vs Grok 3
DeepSeek-V4 Flash is 95.3% cheaper for input tokens ($0.14 vs. $3.00 per 1M tokens) and $0.56 vs. $15.00 for output tokens (21.4x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (700ms faster than Grok 3). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 58.5% for Grok 3.
DeepSeek-V4 Flash vs xAI Grok 4.6
DeepSeek-V4 Flash is 93.0% cheaper for input tokens ($0.14 vs. $2.00 per 1M tokens) and $0.56 vs. $6.00 for output tokens (14.3x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (130ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.
DeepSeek-V4 Flash vs Llama 3.3 70B Instruct
DeepSeek-V4 Flash is 22.2% cheaper for input tokens ($0.14 vs. $0.18 per 1M tokens) and $0.56 vs. $0.40 for output tokens (1.3x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (270ms faster than Llama 3.3 70B Instruct). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 38.8% for Llama 3.3 70B Instruct.
DeepSeek-V4 Flash vs Mistral Large 2
DeepSeek-V4 Flash is 93.0% cheaper for input tokens ($0.14 vs. $2.00 per 1M tokens) and $0.56 vs. $6.00 for output tokens (14.3x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (400ms faster than Mistral Large 2). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 39.0% for Mistral Large 2.
DeepSeek-V4 Flash vs OpenAI o1
DeepSeek-V4 Flash is 99.1% cheaper for input tokens ($0.14 vs. $15.00 per 1M tokens) and $0.56 vs. $60.00 for output tokens (107.1x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (700ms faster than OpenAI o1). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 48.9% for OpenAI o1.
DeepSeek-V4 Flash vs o3-mini
DeepSeek-V4 Flash is 87.3% cheaper for input tokens ($0.14 vs. $1.10 per 1M tokens) and $0.56 vs. $4.40 for output tokens (7.9x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (1050ms faster than o3-mini). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 49.3% for o3-mini.
DeepSeek-V4 Flash vs Microsoft Phi-4 (14B)
Microsoft Phi-4 (14B) is 14.3% cheaper for input tokens ($0.12 vs. $0.14 per 1M tokens) and $0.36 vs. $0.56 for output tokens (1.2x cost difference). In terms of operational performance, Microsoft Phi-4 (14B) delivers faster response latency with 110ms TTFT (40ms faster than DeepSeek-V4 Flash). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 42.1% for Microsoft Phi-4 (14B).
DeepSeek-V4 Flash vs Qwen 2.5 72B Instruct
DeepSeek-V4 Flash is 60.0% cheaper for input tokens ($0.14 vs. $0.35 per 1M tokens) and $0.56 vs. $0.40 for output tokens (2.5x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (270ms faster than Qwen 2.5 72B Instruct). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 44.0% for Qwen 2.5 72B Instruct.
DeepSeek-V4 Flash vs Qwen 2.5 Max
DeepSeek-V4 Flash is 50.0% cheaper for input tokens ($0.14 vs. $0.28 per 1M tokens) and $0.56 vs. $0.84 for output tokens (2.0x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (330ms faster than Qwen 2.5 Max). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 44.2% for Qwen 2.5 Max.
DeepSeek-V4 Flash vs Qwen 3.8 Flash Next
Qwen 3.8 Flash Next is 14.3% cheaper for input tokens ($0.12 vs. $0.14 per 1M tokens) and $0.48 vs. $0.56 for output tokens (1.2x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 120ms TTFT (30ms faster than DeepSeek-V4 Flash). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 56.8% for Qwen 3.8 Flash Next.
Frequently Asked Questions & Query Fan-Out
How much does DeepSeek-V4 Flash cost per 1M tokens?
DeepSeek-V4 Flash costs $0.14 per million prompt (input) tokens and $0.56 per million completion (output) tokens.
What is the context window for DeepSeek-V4 Flash?
DeepSeek-V4 Flash supports a maximum context window of 1,024,000 tokens, with a maximum single-generation output of 32,768 tokens.
What are the primary use cases for DeepSeek-V4 Flash?
Ultra-low-cost open-weights reasoning and coding with 1M context at unprecedented token generation speeds.