What are the token costs and operational benchmarks for Gemini 2.5 Pro?

Gemini 2.5 Pro is priced at $1.25 per million input tokens and $10.00 per million output tokens. It features a 1,024k token context window, an average response latency of 700ms TTFT, and achieves 63.8% on SWE-bench Verified and 85.2% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $1.25
Output / 1M $10.00
Context Limit 1,024k
TTFT Latency 700ms
SWE-bench 63.8%
Throughput 80 t/s

Architectural Overview

Frontier foundation model developed by Google featuring Native Multimodal MoE with Thinking architecture.

Optimal Production Use Cases

Million-token multimodal analysis spanning hours of video, audio streams, massive document archives, and complex codebase ingestion.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens with prompt caching economics.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

Gemini 2.5 Pro $37.50

$1.25/1M in · $10.00/1M out

Prompt Cache: $0.313/1M read · $1.25/1M write

Qwen 3.8 Flash Next $1.88

$0.10/1M in · $0.35/1M out

Prompt Cache: $0.020/1M read · $0.10/1M write

Projected Monthly Cost Reduction
$35.63 / mo
(95.0% lower cost with Qwen 3.8 Flash Next)

Head-to-Head Comparisons Involving Gemini 2.5 Pro

Versus Comparison →

Gemini 2.5 Pro vs Claude 3.5 Haiku

Claude 3.5 Haiku is 36.0% cheaper for input tokens ($0.80 vs. $1.25 per 1M tokens) and $4.00 vs. $10.00 for output tokens (1.6x cost difference). In terms of operational performance, Claude 3.5 Haiku delivers faster response latency with 190ms TTFT (160ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 40.6% for Claude 3.5 Haiku.

Versus Comparison →

Gemini 2.5 Pro vs Claude 3.5 Sonnet

Gemini 2.5 Pro is 58.3% cheaper for input tokens ($1.25 vs. $3.00 per 1M tokens) and $10.00 vs. $15.00 for output tokens (2.4x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 310ms TTFT (40ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 49.0% for Claude 3.5 Sonnet.

Versus Comparison →

Gemini 2.5 Pro vs Claude 3.7 Sonnet

**Gemini 2.5 Pro delivers substantial cost savings and speed, while Claude 3.7 Sonnet leads in software engineering capability.** Gemini is 58.3% cheaper on inputs at $1.25 versus $3.00 per 1M tokens, with $10.00 versus $15.00 outputs and a 700ms TTFT (80ms faster). Conversely, Claude outpaces Gemini on SWE-bench, achieving 70.3% compared to 63.8%.

Versus Comparison →

Gemini 2.5 Pro vs DeepSeek R1

DeepSeek-R1 is 56.0% cheaper for input tokens ($0.55 vs. $1.25 per 1M tokens) and $2.19 vs. $10.00 for output tokens (2.3x cost difference). In terms of operational performance, Gemini 2.5 Pro delivers faster response latency with 350ms TTFT (1450ms faster than DeepSeek-R1). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison →

Gemini 2.5 Pro vs DeepSeek V3

**DeepSeek V3 delivers massive cost efficiency, but Gemini 2.5 Pro leads in advanced capabilities.** Gemini outperforms in coding with a 63.8% SWE-bench score versus DeepSeek’s 42.0%. Conversely, DeepSeek V3 is 88.8% cheaper on input ($0.14 versus $1.25 per 1M tokens), $0.28 versus $10.00 for output, and 290ms faster at 410ms TTFT.

Versus Comparison →

Gemini 2.5 Pro vs DeepSeek-V4 Flash Vision Exp

**DeepSeek-V4 Flash Vision Exp delivers unmatched cost efficiency and performance over Gemini 2.5 Pro.** It is 88.8% cheaper for input tokens ($0.14 vs. $1.25/M) and outputs cost $0.56 vs. $10.00/M. DeepSeek also achieves a 140ms TTFT (560ms faster) and slightly edges out coding benchmarks with 64.2% versus 63.8% on SWE-bench.

Versus Comparison →

Gemini 2.5 Pro vs Gemini 2.0 Flash

Gemini 2.0 Flash is 92.0% cheaper for input tokens ($0.10 vs. $1.25 per 1M tokens) and $0.40 vs. $10.00 for output tokens (12.5x cost difference). In terms of operational performance, Gemini 2.0 Flash delivers faster response latency with 220ms TTFT (130ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 28.0% for Gemini 2.0 Flash.

Versus Comparison →

Gemini 2.5 Pro vs Gemini 2.5 Flash

Gemini 2.5 Flash is 76.0% cheaper for input tokens ($0.30 vs. $1.25 per 1M tokens) and $2.50 vs. $10.00 for output tokens (4.2x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (240ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 58.0% for Gemini 2.5 Flash.

Versus Comparison →

Gemini 2.5 Pro vs GLM-5.3

GLM-5.3 is 72.0% cheaper for input tokens ($0.35 vs. $1.25 per 1M tokens) and $1.00 vs. $10.00 for output tokens (3.6x cost difference). In terms of operational performance, GLM-5.3 delivers faster response latency with 160ms TTFT (190ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 56.0% for GLM-5.3.

Versus Comparison →

Gemini 2.5 Pro vs GLM-5.3-Flash

GLM-5.3-Flash is 90.4% cheaper for input tokens ($0.12 vs. $1.25 per 1M tokens) and $0.40 vs. $10.00 for output tokens (10.4x cost difference). In terms of operational performance, GLM-5.3-Flash delivers faster response latency with 90ms TTFT (260ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 52.8% for GLM-5.3-Flash.

Versus Comparison →

Gemini 2.5 Pro vs GPT-4.5

**Gemini 2.5 Pro drastically outperforms GPT-4.5 on cost-efficiency and coding capabilities.** It is 98.3% cheaper for input ($1.25 vs. $75.00/1M tokens) and 60.0x cheaper for output ($10.00 vs. $150.00/1M tokens). Gemini also leads in speed with a 700ms TTFT (150ms faster) and dominates coding at 63.8% versus 38.0% on SWE-bench.

Versus Comparison →

Gemini 2.5 Pro vs GPT-4o

Gemini 2.5 Pro is 50.0% cheaper for input tokens ($1.25 vs. $2.50 per 1M tokens) and $10.00 vs. $10.00 for output tokens (2.0x cost difference). In terms of operational performance, GPT-4o delivers faster response latency with 280ms TTFT (70ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 48.0% for GPT-4o.

Versus Comparison →

Gemini 2.5 Pro vs GPT-4o Mini

GPT-4o Mini is 88.0% cheaper for input tokens ($0.15 vs. $1.25 per 1M tokens) and $0.60 vs. $10.00 for output tokens (8.3x cost difference). In terms of operational performance, GPT-4o Mini delivers faster response latency with 220ms TTFT (130ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 41.0% for GPT-4o Mini.

Versus Comparison →

Gemini 2.5 Pro vs Grok 3

**Gemini 2.5 Pro provides a massive cost advantage over Grok 3, despite Grok leading in speed and coding.** Gemini is 58.3% cheaper for inputs ($1.25 vs. $3.00 per 1M tokens; 2.4x difference) and outputs ($10.00 vs. $15.00). However, Grok 3 delivers faster 620ms TTFT (80ms quicker) and leads SWE-bench at 65.8% versus 63.8%.

Versus Comparison →

Gemini 2.5 Pro vs Grok 3 Mini

Grok 3 Mini is 76.0% cheaper for input tokens ($0.30 vs. $1.25 per 1M tokens) and $1.20 vs. $10.00 for output tokens (4.2x cost difference). In terms of operational performance, Grok 3 Mini delivers faster response latency with 150ms TTFT (200ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 54.0% for Grok 3 Mini.

Versus Comparison →

Gemini 2.5 Pro vs Llama 3.3 70B Instruct

**Llama 3.3 70B Instruct offers unbeatable cost efficiency and speed, while Gemini 2.5 Pro leads complex capability.** Llama 3.3 is 90.4% cheaper for inputs ($0.12 vs. $1.25/1M) and outputs ($0.30 vs. $10.00), running at 450ms TTFT (250ms faster). Conversely, Gemini 2.5 Pro dominates coding, achieving 63.8% on SWE-bench versus 42.5%.

Versus Comparison →

Gemini 2.5 Pro vs Mistral Large 2

**Gemini 2.5 Pro dominates complex coding capability**, achieving 63.8% on SWE-bench versus Mistral Large 2's 40.0%. Gemini is 37.5% cheaper for inputs at $1.25/1M tokens, though Mistral offers cheaper outputs ($6.00 vs. $10.00/1M). However, Mistral Large 2 provides faster initial responsiveness with a 520ms TTFT, leading Gemini by 180ms.

Versus Comparison →

Gemini 2.5 Pro vs Mistral Large 2 (2411)

**Gemini 2.5 Pro dominates complex coding performance with a 63.8% SWE-bench score versus Mistral Large 2411’s 37.8%.** Gemini provides 37.5% cheaper input pricing at $1.25 compared to $2.00 per 1M tokens, but higher output costs ($10.00 vs. $6.00). Conversely, Mistral delivers faster responsiveness, clocking a 360ms TTFT that is 340ms faster than Gemini.

Versus Comparison →

Gemini 2.5 Pro vs o1

**Gemini 2.5 Pro decisively outperforms o1 across efficiency and capability benchmarks.** It leads SWE-bench at 65.0% versus 48.9% for o1, delivering a 350ms TTFT that is 2250ms faster. Financially, Gemini 2.5 Pro is 91.7% cheaper for input tokens at $1.25 versus $15.00 per million, and $10.00 versus $60.00 for output tokens.

Versus Comparison →

Gemini 2.5 Pro vs OpenAI o3-mini

**Gemini 2.5 Pro dominates in capability and speed, despite OpenAI o3-mini offering lower operational costs.** Gemini outperforms in coding with 63.8% on SWE-bench versus 49.3%, alongside a faster 700ms TTFT (500ms advantage). Conversely, o3-mini is 12.0% cheaper on input ($1.10 vs. $1.25/1M) and more economical on output ($4.40 vs. $10.00/1M).

Versus Comparison →

Gemini 2.5 Pro vs Qwen 3.8 27B

Qwen 3.8 27B is 84.0% cheaper for input tokens ($0.20 vs. $1.25 per 1M tokens) and $0.60 vs. $10.00 for output tokens (6.2x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (230ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 58.4% for Qwen 3.8 27B.

Versus Comparison →

Gemini 2.5 Pro vs Qwen 3.8 Flash Next

Qwen 3.8 Flash Next is 92.0% cheaper for input tokens ($0.10 vs. $1.25 per 1M tokens) and $0.35 vs. $10.00 for output tokens (12.5x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 95ms TTFT (255ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 61.0% for Qwen 3.8 Flash Next.

Frequently Asked Questions & Query Fan-Out

How much does Gemini 2.5 Pro cost per 1M tokens?

Gemini 2.5 Pro costs $1.25 per million prompt (input) tokens and $10.00 per million completion (output) tokens.

What is the context window for Gemini 2.5 Pro?

Gemini 2.5 Pro supports a maximum context window of 1,024,000 tokens, with a maximum single-generation output of 65,536 tokens.

What are the primary use cases for Gemini 2.5 Pro?

Million-token multimodal analysis spanning hours of video, audio streams, massive document archives, and complex codebase ingestion.