What are the token costs and operational benchmarks for Mistral Large 2?

Mistral Large 2 is priced at $2.00 per million input tokens and $6.00 per million output tokens. It features a 131.072k token context window, an average response latency of 520ms TTFT, and achieves 40% on SWE-bench Verified and 74.5% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $2.00
Output / 1M $6.00
Context Limit 131.072k
TTFT Latency 520ms
SWE-bench 40%
Throughput 85 t/s

Architectural Overview

Frontier foundation model developed by Mistral AI featuring Dense Transformer architecture.

Optimal Production Use Cases

European data-sovereign enterprise workflows, multilingual translation across European languages, and precise JSON tool execution.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens with prompt caching economics.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

Mistral Large 2 $35.00

$2.00/1M in · $6.00/1M out

Prompt Cache: $0.500/1M read · $2.00/1M write

Qwen 3.8 Flash Next $1.88

$0.10/1M in · $0.35/1M out

Prompt Cache: $0.020/1M read · $0.10/1M write

Projected Monthly Cost Reduction
$33.13 / mo
(94.6% lower cost with Qwen 3.8 Flash Next)

Head-to-Head Comparisons Involving Mistral Large 2

Versus Comparison →

Mistral Large 2 vs Claude 3.5 Haiku

**Claude 3.5 Haiku offers superior cost-efficiency and speed without sacrificing frontier coding performance.** It is 60.0% cheaper on inputs ($0.80 vs. $2.00 per 1M tokens) and costs $4.00 vs. $6.00 for outputs. Furthermore, Claude provides a 190ms TTFT (330ms faster) while edging out Mistral Large 2 on SWE-bench (40.6% vs. 40.0%).

Versus Comparison →

Mistral Large 2 vs Claude 3.5 Sonnet

**Claude 3.5 Sonnet dominates in capability and speed, while Mistral Large 2 provides superior cost efficiency.** Sonnet leads SWE-bench at 49.0% versus 40.0% and yields a faster 310ms TTFT (210ms lower latency). However, Mistral Large 2 is 33.3% cheaper on input ($2.00 vs. $3.00/1M) and costs $6.00 versus $15.00/1M for output tokens.

Versus Comparison →

Mistral Large 2 vs Claude 3.7 Sonnet

**Claude 3.7 Sonnet delivers superior coding capability, while Mistral Large 2 offers significant cost and latency advantages.** Claude 3.7 dominates SWE-bench at 70.3% versus Mistral's 40.0%. However, Mistral Large 2 is 33.3% cheaper on inputs ($2.00 vs. $3.00/1M), cuts outputs ($6.00 vs. $15.00/1M), and responds 130ms faster with a 520ms TTFT.

Versus Comparison →

Mistral Large 2 vs DeepSeek R1

**DeepSeek-R1 delivers superior coding capability at dramatically lower cost.** DeepSeek-R1 scores 52.0% on SWE-bench versus 40.0% for Mistral Large 2, while costing 72.5% less for inputs ($0.55 vs. $2.00/1M) and $2.19 vs. $6.00 for outputs. Conversely, Mistral Large 2 leads in responsiveness with a 520ms TTFT, 580ms faster than DeepSeek-R1.

Versus Comparison →

Mistral Large 2 vs DeepSeek V3

**DeepSeek V3 delivers massive price-performance advantages over Mistral Large 2, highlighted by a 14.3x lower overall cost.** DeepSeek V3 is 93.0% cheaper for inputs ($0.14 vs. $2.00/1M) and outputs ($0.28 vs. $6.00/1M). It also delivers faster latency at 410ms TTFT (110ms faster) and leads coding benchmarks with 42.0% versus 40.0% on SWE-bench.

Versus Comparison →

Mistral Large 2 vs DeepSeek-V4 Flash Vision Exp

**DeepSeek-V4 Flash Vision Exp decisively beats Mistral Large 2 on both cost and capability.** DeepSeek is 93.0% cheaper, pricing input tokens at $0.14 (versus $2.00) and output at $0.56 (versus $6.00) per 1M tokens. It responds faster with a 140ms TTFT (380ms advantage) while leading SWE-bench coding benchmarks at 64.2% versus 40.0%.

Versus Comparison →

Mistral Large 2 vs Gemini 2.0 Flash

**Gemini 2.0 Flash provides unmatched cost efficiency and speed over Mistral Large 2.** It is 95.0% cheaper for input ($0.10 vs. $2.00 per 1M tokens) and output ($0.40 vs. $6.00). Gemini also responds faster with a 220ms TTFT (300ms advantage) while outperforming Mistral on SWE-bench coding benchmarks at 45.2% versus 40.0%.

Versus Comparison →

Mistral Large 2 vs Gemini 2.5 Flash

**Gemini 2.5 Flash provides vastly superior economics and coding performance over Mistral Large 2.** Gemini is 85.0% cheaper for input ($0.30 vs. $2.00 per 1M) with lower output costs ($2.50 vs. $6.00). It also leads coding capability at 58.0% SWE-bench versus 40.0%, while delivering faster responsiveness at 110ms TTFT (410ms faster).

Versus Comparison →

Mistral Large 2 vs Gemini 2.5 Pro

**Gemini 2.5 Pro dominates complex coding capability**, achieving 63.8% on SWE-bench versus Mistral Large 2's 40.0%. Gemini is 37.5% cheaper for inputs at $1.25/1M tokens, though Mistral offers cheaper outputs ($6.00 vs. $10.00/1M). However, Mistral Large 2 provides faster initial responsiveness with a 520ms TTFT, leading Gemini by 180ms.

Versus Comparison →

Mistral Large 2 vs GLM-5.3-Flash

**GLM-5.3-Flash delivers superior performance at a fraction of the cost, making it the clear winner over Mistral Large 2.** It is 94.0% cheaper, pricing input tokens at $0.12 versus $2.00 and output at $0.40 versus $6.00 per 1M tokens. GLM-5.3-Flash also cuts latency by 430ms to 90ms TTFT while leading SWE-bench 52.8% to 40.0%.

Versus Comparison →

Mistral Large 2 vs GLM-5.3

**GLM-5.3 significantly outperforms Mistral Large 2 across cost efficiency, latency, and coding performance.** GLM-5.3 is 82.5% cheaper on inputs ($0.35 vs. $2.00 per 1M) and 5.7x cheaper on outputs ($1.00 vs. $6.00). It also delivers faster 160ms TTFT (360ms lower latency) while leading SWE-bench coding benchmarks at 56.0% compared to 40.0%.

Versus Comparison →

Mistral Large 2 vs GPT-4.5

**Mistral Large 2 delivers vastly superior cost efficiency and faster execution than GPT-4.5.** It is 97.3% cheaper at $2.00 input and $6.00 output per 1M tokens versus $75.00 and $150.00 (a 37.5x cost difference). Mistral Large 2 also achieves a 520ms TTFT (330ms faster

Versus Comparison →

Mistral Large 2 vs GPT-4o Mini

**GPT-4o Mini delivers massive cost efficiency and speed advantages over Mistral Large 2.** It is 92.5% cheaper for inputs ($0.15 vs. $2.00 per 1M) and $0.60 vs. $6.00 for outputs. Additionally, GPT-4o Mini reduces TTFT latency to 220ms (300ms faster) while slightly outperforming Mistral Large 2 on SWE-bench coding benchmarks (41.0% vs. 40.0%).

Versus Comparison →

Mistral Large 2 vs GPT-4o

**GPT-4o outperforms Mistral Large 2 in speed and capability, while Mistral Large 2 provides better cost efficiency.** Mistral Large 2 is 20.0% cheaper for input at $2.00 versus $2.50 per 1M tokens ($6.00 vs. $10.00 output). However, GPT-4o achieves a faster 280ms TTFT (240ms lower) and leads SWE-bench at 48.0% over 40.0%.

Versus Comparison →

Mistral Large 2 vs Grok 3 Mini

**Grok 3 Mini dominates Mistral Large 2 with superior coding capability and drastically lower costs.** It is 85.0% cheaper for input tokens ($0.30 vs. $2.00 per 1M) and $1.20 vs. $6.00 for output. Additionally, Grok 3 Mini achieves a faster 150ms TTFT (370ms advantage) while outperforming on SWE-bench at 54.0% compared to Mistral Large 2's 40.0%.

Versus Comparison →

Mistral Large 2 vs Grok 3

**Grok 3 delivers superior coding performance, but Mistral Large 2 provides unmatched cost efficiency.** Grok 3 dominates with 58.0% on SWE-bench versus 40.0%. Conversely, Mistral Large 2 is 33.3% cheaper for inputs ($2.00 vs. $3.00/1M), cuts output costs to $6.00 vs. $15.00, and clocks a faster 520ms TTFT (230ms advantage).

Versus Comparison →

Mistral Large 2 vs Llama 3.3 70B Instruct

**Llama 3.3 70B Instruct delivers vastly superior economic value and speed over Mistral Large 2.** It is 94.0% cheaper for inputs ($0.12 vs. $2.00/1M tokens) and outputs ($0.30 vs. $6.00/1M). Additionally, Llama 3.3 responds faster with a 450ms TTFT (70ms lead) while outperforming in coding at 42.5% versus 40.0% SWE-bench.

Versus Comparison →

Mistral Large 2 vs Mistral Large 2 (2411)

**Mistral Large 2 delivers superior coding performance and responsiveness at identical pricing.** At $2.00 per 1M input tokens for both models, Mistral Large 2 leads software engineering tasks with a 40.0% SWE-bench score versus 30.1% for Mistral Large 2 (2411). It also reduces latency, achieving a 520ms TTFT, exactly 60ms faster than 2411.

Versus Comparison →

Mistral Large 2 vs o1

**Mistral Large 2 offers massive cost and latency advantages, while o1 delivers superior reasoning performance.** Mistral Large 2 is 86.7% cheaper on inputs ($2.00 vs. $15.00/1M) and $6.00 vs. $60.00 on outputs, while operating 2080ms faster at 520ms TTFT. Conversely, o1 leads software engineering benchmarks, scoring 48.9% on SWE-bench compared to Mistral’s 40.0%.

Versus Comparison →

Mistral Large 2 vs OpenAI o3-mini

**OpenAI o3-mini delivers superior coding performance and substantial cost savings over Mistral Large 2.** It leads coding benchmarks at 49.3% SWE-bench versus 40.0% while cutting costs by 45.0% on inputs ($1.10 vs. $2.00/1M) and $4.40 vs. $6.00 for outputs. Conversely, Mistral Large 2 wins on speed with a 520ms TTFT, running 680ms faster.

Versus Comparison →

Mistral Large 2 vs Qwen 3.8 27B

**Qwen 3.8 27B delivers superior capability at a 90.0% lower cost than Mistral Large 2.** Qwen costs just $0.20 input and $0.60 output per 1M tokens versus Mistral's $2.00 and $6.00. Furthermore, Qwen achieves a faster 120ms TTFT (a 400ms advantage) and leads coding with a 58.4% SWE-bench score over Mistral’s 40.0%.

Versus Comparison →

Mistral Large 2 vs Qwen 3.8 Flash Next

**Qwen 3.8 Flash Next delivers superior coding performance at a fraction of the cost.** It is 95.0% cheaper on input tokens ($0.10 vs. $2.00/1M) and 20.0x cheaper on output ($0.35 vs. $6.00/1M). Qwen also dominates latency with a 95ms TTFT (425ms faster) while outperforming Mistral Large 2 on SWE-bench (61.0% vs. 40.0%).

Frequently Asked Questions & Query Fan-Out

How much does Mistral Large 2 cost per 1M tokens?

Mistral Large 2 costs $2.00 per million prompt (input) tokens and $6.00 per million completion (output) tokens.

What is the context window for Mistral Large 2?

Mistral Large 2 supports a maximum context window of 131,072 tokens, with a maximum single-generation output of 8,192 tokens.

What are the primary use cases for Mistral Large 2?

European data-sovereign enterprise workflows, multilingual translation across European languages, and precise JSON tool execution.