What are the token costs and operational benchmarks for Gemini 2.5 Flash?

Gemini 2.5 Flash is priced at $0.30 per million input tokens and $2.50 per million output tokens. It features a 1,048.576k token context window, an average response latency of 110ms TTFT, and achieves 58% on SWE-bench Verified and 82% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $0.30
Output / 1M $2.50
Context Limit 1,048.576k
TTFT Latency 110ms
SWE-bench 58%
Throughput 180 t/s

Architectural Overview

Frontier foundation model developed by Google featuring Hybrid Reasoning Omni Transformer architecture.

Optimal Production Use Cases

Hybrid reasoning model delivering unmatched price-performance across 1M context with configurable thinking budgets.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens with prompt caching economics.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

Gemini 2.5 Flash $9.25

$0.30/1M in · $2.50/1M out

Prompt Cache: $0.030/1M read · $0.30/1M write

Qwen 3.8 Flash Next $1.88

$0.10/1M in · $0.35/1M out

Prompt Cache: $0.020/1M read · $0.10/1M write

Projected Monthly Cost Reduction
$7.38 / mo
(79.7% lower cost with Qwen 3.8 Flash Next)

Head-to-Head Comparisons Involving Gemini 2.5 Flash

Versus Comparison →

Gemini 2.5 Flash vs Claude 3.5 Haiku

Gemini 2.5 Flash is 62.5% cheaper for input tokens ($0.30 vs. $0.80 per 1M tokens) and $2.50 vs. $4.00 for output tokens (2.7x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (80ms faster than Claude 3.5 Haiku). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 40.6% for Claude 3.5 Haiku.

Versus Comparison →

Gemini 2.5 Flash vs Claude 3.5 Sonnet

Gemini 2.5 Flash is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $2.50 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (200ms faster than Claude 3.5 Sonnet). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 49.0% for Claude 3.5 Sonnet.

Versus Comparison →

Gemini 2.5 Flash vs Claude 3.7 Sonnet

Gemini 2.5 Flash is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $2.50 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (540ms faster than Claude 3.7 Sonnet). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 58.0% for Gemini 2.5 Flash.

Versus Comparison →

Gemini 2.5 Flash vs DeepSeek R1

Gemini 2.5 Flash is 45.5% cheaper for input tokens ($0.30 vs. $0.55 per 1M tokens) and $2.50 vs. $2.19 for output tokens (1.8x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (1690ms faster than DeepSeek-R1). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison →

Gemini 2.5 Flash vs DeepSeek V3

**DeepSeek V3 delivers massive cost savings, whereas Gemini 2.5 Flash prioritizes low latency and coding performance.** DeepSeek V3 is 53.3% cheaper for inputs ($0.14 vs. $0.30 per 1M) and $0.28 versus $2.50 for outputs. Conversely, Gemini 2.5 Flash achieves a 110ms TTFT (300ms faster) and leads coding benchmarks with 58.0% over 42.0% on SWE-bench.

Versus Comparison →

Gemini 2.5 Flash vs DeepSeek-V4 Flash Vision Exp

DeepSeek-V4 Flash Vision Exp is 53.3% cheaper for input tokens ($0.14 vs. $0.30 per 1M tokens) and $0.56 vs. $2.50 for output tokens (2.1x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (30ms faster than DeepSeek-V4 Flash Vision Exp). DeepSeek-V4 Flash Vision Exp leads coding benchmarks at 64.2% SWE-bench vs. 58.0% for Gemini 2.5 Flash.

Versus Comparison →

Gemini 2.5 Flash vs Gemini 2.0 Flash

Gemini 2.0 Flash is 66.7% cheaper for input tokens ($0.10 vs. $0.30 per 1M tokens) and $0.40 vs. $2.50 for output tokens (3.0x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (110ms faster than Gemini 2.0 Flash). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 28.0% for Gemini 2.0 Flash.

Versus Comparison →

Gemini 2.5 Flash vs Gemini 2.5 Pro

Gemini 2.5 Flash is 76.0% cheaper for input tokens ($0.30 vs. $1.25 per 1M tokens) and $2.50 vs. $10.00 for output tokens (4.2x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (240ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 58.0% for Gemini 2.5 Flash.

Versus Comparison →

Gemini 2.5 Flash vs GLM-5.3

Gemini 2.5 Flash is 14.3% cheaper for input tokens ($0.30 vs. $0.35 per 1M tokens) and $2.50 vs. $1.00 for output tokens (1.2x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (50ms faster than GLM-5.3). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 56.0% for GLM-5.3.

Versus Comparison →

Gemini 2.5 Flash vs GLM-5.3-Flash

GLM-5.3-Flash is 60.0% cheaper for input tokens ($0.12 vs. $0.30 per 1M tokens) and $0.40 vs. $2.50 for output tokens (2.5x cost difference). In terms of operational performance, GLM-5.3-Flash delivers faster response latency with 90ms TTFT (20ms faster than Gemini 2.5 Flash). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 52.8% for GLM-5.3-Flash.

Versus Comparison →

Gemini 2.5 Flash vs GPT-4.5

**Gemini 2.5 Flash delivers a massive price-performance advantage over GPT-4.5.** It is 99.6% cheaper for input tokens at $0.30 versus $75.00 per 1M tokens ($2.50 vs. $150.00 output). Additionally, Gemini 2.5 Flash is 740ms faster with a 110ms TTFT and leads coding benchmarks with 58.0% on SWE-bench compared to GPT-4.5's 38.0%.

Versus Comparison →

Gemini 2.5 Flash vs GPT-4o

Gemini 2.5 Flash is 88.0% cheaper for input tokens ($0.30 vs. $2.50 per 1M tokens) and $2.50 vs. $10.00 for output tokens (8.3x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (170ms faster than GPT-4o). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 48.0% for GPT-4o.

Versus Comparison →

Gemini 2.5 Flash vs GPT-4o Mini

GPT-4o Mini is 50.0% cheaper for input tokens ($0.15 vs. $0.30 per 1M tokens) and $0.60 vs. $2.50 for output tokens (2.0x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (110ms faster than GPT-4o Mini). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 41.0% for GPT-4o Mini.

Versus Comparison →

Gemini 2.5 Flash vs Grok 3

**Gemini 2.5 Flash provides massive cost and latency advantages over Grok 3.** It is 90.0% cheaper for inputs ($0.30 versus $3.00 per 1M tokens) and outputs ($2.50 versus $15.00), clocking a 110ms TTFT (510ms faster). Conversely, Grok 3 justifies its premium for advanced tasks, leading SWE-bench coding benchmarks at 65.8% compared to Gemini's 58.0%.

Versus Comparison →

Gemini 2.5 Flash vs Grok 3 Mini

Both models share identical input pricing at $0.30 per 1M tokens. In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (40ms faster than Grok 3 Mini). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 54.0% for Grok 3 Mini.

Versus Comparison →

Gemini 2.5 Flash vs Llama 3.3 70B Instruct

**Gemini 2.5 Flash delivers superior capability and speed, while Llama 3.3 70B Instruct maximizes cost efficiency.** Gemini leads coding benchmarks at 58.0% SWE-bench versus 42.5% and responds 340ms faster with 110ms TTFT. Conversely, Llama is 60.0% cheaper for input tokens ($0.12 vs. $0.30/1M) and drastically cheaper on output ($0.30 vs. $2.50/1M).

Versus Comparison →

Gemini 2.5 Flash vs Mistral Large 2

**Gemini 2.5 Flash provides vastly superior economics and coding performance over Mistral Large 2.** Gemini is 85.0% cheaper for input ($0.30 vs. $2.00 per 1M) with lower output costs ($2.50 vs. $6.00). It also leads coding capability at 58.0% SWE-bench versus 40.0%, while delivering faster responsiveness at 110ms TTFT (410ms faster).

Versus Comparison →

Gemini 2.5 Flash vs Mistral Large 2 (2411)

Gemini 2.5 Flash is 85.0% cheaper for input tokens ($0.30 vs. $2.00 per 1M tokens) and $2.50 vs. $6.00 for output tokens (6.7x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (340ms faster than Mistral Large 2411). Gemini 2.5 Flash leads coding benchmarks at 58.0% SWE-bench vs. 35.7% for Mistral Large 2411.

Versus Comparison →

Gemini 2.5 Flash vs o1

**Gemini 2.5 Flash delivers superior coding performance at a fraction of o1's cost.** It is 98.0% cheaper for input tokens ($0.30 vs. $15.00/1M) and lower on outputs ($2.50 vs. $60.00/1M). Furthermore, Gemini achieves a 110ms TTFT (2490ms faster) while leading SWE-bench scoring at 58.0% compared to o1's 48.9%.

Versus Comparison →

Gemini 2.5 Flash vs OpenAI o3-mini

**Gemini 2.5 Flash delivers superior overall value, beating OpenAI o3-mini across cost, latency, and coding benchmarks.** Gemini 2.5 Flash is 72.7% cheaper for inputs ($0.30 vs. $1.10/1M tokens) and lower on outputs ($2.50 vs. $4.40). It responds 1090ms faster with a 110ms TTFT and leads SWE-bench (58.0% vs. 49.3%).

Versus Comparison →

Gemini 2.5 Flash vs Qwen 3.8 27B

Qwen 3.8 27B is 33.3% cheaper for input tokens ($0.20 vs. $0.30 per 1M tokens) and $0.60 vs. $2.50 for output tokens (1.5x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (10ms faster than Qwen 3.8 27B). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 58.0% for Gemini 2.5 Flash.

Versus Comparison →

Gemini 2.5 Flash vs Qwen 3.8 Flash Next

Qwen 3.8 Flash Next is 66.7% cheaper for input tokens ($0.10 vs. $0.30 per 1M tokens) and $0.35 vs. $2.50 for output tokens (3.0x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 95ms TTFT (15ms faster than Gemini 2.5 Flash). Qwen 3.8 Flash Next leads coding benchmarks at 61.0% SWE-bench vs. 58.0% for Gemini 2.5 Flash.

Frequently Asked Questions & Query Fan-Out

How much does Gemini 2.5 Flash cost per 1M tokens?

Gemini 2.5 Flash costs $0.30 per million prompt (input) tokens and $2.50 per million completion (output) tokens.

What is the context window for Gemini 2.5 Flash?

Gemini 2.5 Flash supports a maximum context window of 1,048,576 tokens, with a maximum single-generation output of 65,536 tokens.

What are the primary use cases for Gemini 2.5 Flash?

Hybrid reasoning model delivering unmatched price-performance across 1M context with configurable thinking budgets.