Gemini 2.5 Flash
Proprietary CommercialDeveloped by Google · Released 2025-03-03
What are the token costs and operational benchmarks for Gemini 2.5 Flash?
Architectural Overview
Frontier foundation model developed by Google featuring Hybrid Reasoning Omni Transformer architecture.
Optimal Production Use Cases
Hybrid reasoning model delivering unmatched price-performance across 1M context with configurable thinking budgets.
Interactive Monthly Token Economics & ROI Forecaster
Model your expected production workload across prompt (input) and completion (output) tokens with prompt caching economics.
Estimated Monthly Spend
$0.30/1M in · $2.50/1M out
Prompt Cache: $0.030/1M read · $0.30/1M write
$0.10/1M in · $0.35/1M out
Prompt Cache: $0.020/1M read · $0.10/1M write
Head-to-Head Comparisons Involving Gemini 2.5 Flash
Gemini 2.5 Flash vs Claude 3.5 Haiku
Gemini 2.5 Flash is 62.5% cheaper for input tokens ($0.30 vs. $0.80 per 1M tokens) and $2.50 vs. $4.00 for output tokens (2.7x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (80ms faster than Claude 3.5 Haiku). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 40.6% for Claude 3.5 Haiku.
Gemini 2.5 Flash vs Claude 3.5 Sonnet
Gemini 2.5 Flash is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $2.50 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (200ms faster than Claude 3.5 Sonnet). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 49.0% for Claude 3.5 Sonnet.
Gemini 2.5 Flash vs Claude 3.7 Sonnet
Gemini 2.5 Flash is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $2.50 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (540ms faster than Claude 3.7 Sonnet). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 58.0% for Gemini 2.5 Flash.
Gemini 2.5 Flash vs DeepSeek R1
Gemini 2.5 Flash is 45.5% cheaper for input tokens ($0.30 vs. $0.55 per 1M tokens) and $2.50 vs. $2.19 for output tokens (1.8x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (1690ms faster than DeepSeek-R1). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 49.2% for DeepSeek-R1.
Gemini 2.5 Flash vs DeepSeek V3
**DeepSeek V3 delivers massive cost savings, whereas Gemini 2.5 Flash prioritizes low latency and coding performance.** DeepSeek V3 is 53.3% cheaper for inputs ($0.14 vs. $0.30 per 1M) and $0.28 versus $2.50 for outputs. Conversely, Gemini 2.5 Flash achieves a 110ms TTFT (300ms faster) and leads coding benchmarks with 58.0% over 42.0% on SWE-bench.
Gemini 2.5 Flash vs DeepSeek-V4 Flash Vision Exp
DeepSeek-V4 Flash Vision Exp is 53.3% cheaper for input tokens ($0.14 vs. $0.30 per 1M tokens) and $0.56 vs. $2.50 for output tokens (2.1x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (30ms faster than DeepSeek-V4 Flash Vision Exp). DeepSeek-V4 Flash Vision Exp leads coding benchmarks at 64.2% SWE-bench vs. 58.0% for Gemini 2.5 Flash.
Gemini 2.5 Flash vs Gemini 2.0 Flash
Gemini 2.0 Flash is 66.7% cheaper for input tokens ($0.10 vs. $0.30 per 1M tokens) and $0.40 vs. $2.50 for output tokens (3.0x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (110ms faster than Gemini 2.0 Flash). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 28.0% for Gemini 2.0 Flash.
Gemini 2.5 Flash vs Gemini 2.5 Pro
Gemini 2.5 Flash is 76.0% cheaper for input tokens ($0.30 vs. $1.25 per 1M tokens) and $2.50 vs. $10.00 for output tokens (4.2x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (240ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 58.0% for Gemini 2.5 Flash.
Gemini 2.5 Flash vs GLM-5.3
Gemini 2.5 Flash is 14.3% cheaper for input tokens ($0.30 vs. $0.35 per 1M tokens) and $2.50 vs. $1.00 for output tokens (1.2x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (50ms faster than GLM-5.3). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 56.0% for GLM-5.3.
Gemini 2.5 Flash vs GLM-5.3-Flash
GLM-5.3-Flash is 60.0% cheaper for input tokens ($0.12 vs. $0.30 per 1M tokens) and $0.40 vs. $2.50 for output tokens (2.5x cost difference). In terms of operational performance, GLM-5.3-Flash delivers faster response latency with 90ms TTFT (20ms faster than Gemini 2.5 Flash). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 52.8% for GLM-5.3-Flash.
Gemini 2.5 Flash vs GPT-4.5
**Gemini 2.5 Flash delivers a massive price-performance advantage over GPT-4.5.** It is 99.6% cheaper for input tokens at $0.30 versus $75.00 per 1M tokens ($2.50 vs. $150.00 output). Additionally, Gemini 2.5 Flash is 740ms faster with a 110ms TTFT and leads coding benchmarks with 58.0% on SWE-bench compared to GPT-4.5's 38.0%.
Gemini 2.5 Flash vs GPT-4o
Gemini 2.5 Flash is 88.0% cheaper for input tokens ($0.30 vs. $2.50 per 1M tokens) and $2.50 vs. $10.00 for output tokens (8.3x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (170ms faster than GPT-4o). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 48.0% for GPT-4o.
Gemini 2.5 Flash vs GPT-4o Mini
GPT-4o Mini is 50.0% cheaper for input tokens ($0.15 vs. $0.30 per 1M tokens) and $0.60 vs. $2.50 for output tokens (2.0x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (110ms faster than GPT-4o Mini). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 41.0% for GPT-4o Mini.
Gemini 2.5 Flash vs Grok 3
**Gemini 2.5 Flash provides massive cost and latency advantages over Grok 3.** It is 90.0% cheaper for inputs ($0.30 versus $3.00 per 1M tokens) and outputs ($2.50 versus $15.00), clocking a 110ms TTFT (510ms faster). Conversely, Grok 3 justifies its premium for advanced tasks, leading SWE-bench coding benchmarks at 65.8% compared to Gemini's 58.0%.
Gemini 2.5 Flash vs Grok 3 Mini
Both models share identical input pricing at $0.30 per 1M tokens. In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (40ms faster than Grok 3 Mini). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 54.0% for Grok 3 Mini.
Gemini 2.5 Flash vs Llama 3.3 70B Instruct
**Gemini 2.5 Flash delivers superior capability and speed, while Llama 3.3 70B Instruct maximizes cost efficiency.** Gemini leads coding benchmarks at 58.0% SWE-bench versus 42.5% and responds 340ms faster with 110ms TTFT. Conversely, Llama is 60.0% cheaper for input tokens ($0.12 vs. $0.30/1M) and drastically cheaper on output ($0.30 vs. $2.50/1M).
Gemini 2.5 Flash vs Mistral Large 2
**Gemini 2.5 Flash provides vastly superior economics and coding performance over Mistral Large 2.** Gemini is 85.0% cheaper for input ($0.30 vs. $2.00 per 1M) with lower output costs ($2.50 vs. $6.00). It also leads coding capability at 58.0% SWE-bench versus 40.0%, while delivering faster responsiveness at 110ms TTFT (410ms faster).
Gemini 2.5 Flash vs Mistral Large 2 (2411)
Gemini 2.5 Flash is 85.0% cheaper for input tokens ($0.30 vs. $2.00 per 1M tokens) and $2.50 vs. $6.00 for output tokens (6.7x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (340ms faster than Mistral Large 2411). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 35.7% for Mistral Large 2411.
Gemini 2.5 Flash vs o1
**Gemini 2.5 Flash delivers superior coding performance at a fraction of o1's cost.** It is 98.0% cheaper for input tokens ($0.30 vs. $15.00/1M) and lower on outputs ($2.50 vs. $60.00/1M). Furthermore, Gemini achieves a 110ms TTFT (2490ms faster) while leading SWE-bench scoring at 58.0% compared to o1's 48.9%.
Gemini 2.5 Flash vs OpenAI o3-mini
**Gemini 2.5 Flash delivers superior overall value, beating OpenAI o3-mini across cost, latency, and coding benchmarks.** Gemini 2.5 Flash is 72.7% cheaper for inputs ($0.30 vs. $1.10/1M tokens) and lower on outputs ($2.50 vs. $4.40). It responds 1090ms faster with a 110ms TTFT and leads SWE-bench (58.0% vs. 49.3%).
Gemini 2.5 Flash vs Qwen 3.8 27B
Qwen 3.8 27B is 33.3% cheaper for input tokens ($0.20 vs. $0.30 per 1M tokens) and $0.60 vs. $2.50 for output tokens (1.5x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (10ms faster than Qwen 3.8 27B). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 58.0% for Gemini 2.5 Flash.
Gemini 2.5 Flash vs Qwen 3.8 Flash Next
Qwen 3.8 Flash Next is 66.7% cheaper for input tokens ($0.10 vs. $0.30 per 1M tokens) and $0.35 vs. $2.50 for output tokens (3.0x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 95ms TTFT (15ms faster than Gemini 2.5 Flash). Qwen 3.8 Flash Next leads coding benchmarks at 61.0% SWE-bench vs. 58.0% for Gemini 2.5 Flash.
Frequently Asked Questions & Query Fan-Out
How much does Gemini 2.5 Flash cost per 1M tokens?
Gemini 2.5 Flash costs $0.30 per million prompt (input) tokens and $2.50 per million completion (output) tokens.
What is the context window for Gemini 2.5 Flash?
Gemini 2.5 Flash supports a maximum context window of 1,048,576 tokens, with a maximum single-generation output of 65,536 tokens.
What are the primary use cases for Gemini 2.5 Flash?
Hybrid reasoning model delivering unmatched price-performance across 1M context with configurable thinking budgets.