What are the token costs and operational benchmarks for Codestral 25.01?

Codestral 25.01 is priced at $0.30 per million input tokens and $0.90 per million output tokens. It features a 262.144k token context window, an average response latency of 150ms TTFT, and achieves 44.2% on SWE-bench Verified and 62% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $0.30
Output / 1M $0.90
Context Limit 262.144k
TTFT Latency 150ms
SWE-bench 44.2%
Throughput 135 t/s

Architectural Overview

Frontier foundation model developed by Mistral AI featuring Dense 256k Code Specialist architecture.

Optimal Production Use Cases

256K long-context repository indexing, fill-in-the-middle code completion, and IDE plugin backend.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

Codestral 25.01 $5.45

$0.3/1M in · $0.9/1M out

GPT-5.6 Luna $67.50

$0.18/1M in · $0.72/1M out

Projected Monthly Cost Reduction
$62.05 / mo
(91.9% lower cost)

Head-to-Head Comparisons Involving Codestral 25.01

Versus Comparison

Codestral 25.01 vs Claude 3.5 Haiku

Codestral 25.01 is 62.5% cheaper for input tokens ($0.30 vs. $0.80 per 1M tokens) and $0.90 vs. $4.00 for output tokens (2.7x cost difference). In terms of operational performance, Claude 3.5 Haiku delivers faster response latency with 140ms TTFT (10ms faster than Codestral 25.01). Codestral 25.01 leads coding benchmarks at 44.2% SWE-bench vs. 40.6% for Claude 3.5 Haiku.

Versus Comparison

Codestral 25.01 vs Claude 3.5 Sonnet

Codestral 25.01 is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $0.90 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (170ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Claude 3.7 Sonnet

Codestral 25.01 is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $0.90 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (500ms faster than Claude 3.7 Sonnet). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Claude Opus 5

Codestral 25.01 is 94.0% cheaper for input tokens ($0.30 vs. $5.00 per 1M tokens) and $0.90 vs. $25.00 for output tokens (16.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (190ms faster than Claude Opus 5). Claude Opus 5 leads coding benchmarks at 82.4% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Composer 2.5

Codestral 25.01 is 83.3% cheaper for input tokens ($0.30 vs. $1.80 per 1M tokens) and $0.90 vs. $7.20 for output tokens (6.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (30ms faster than Composer 2.5). Composer 2.5 leads coding benchmarks at 74.6% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs DeepSeek-R1

Codestral 25.01 is 45.5% cheaper for input tokens ($0.30 vs. $0.55 per 1M tokens) and $0.90 vs. $2.19 for output tokens (1.8x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (1650ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs DeepSeek-V3

DeepSeek-V3 is 53.3% cheaper for input tokens ($0.14 vs. $0.30 per 1M tokens) and $0.28 vs. $0.90 for output tokens (2.1x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (190ms faster than DeepSeek-V3). Codestral 25.01 leads coding benchmarks at 44.2% SWE-bench vs. 42.0% for DeepSeek-V3.

Versus Comparison

Codestral 25.01 vs DeepSeek-V4 Flash

DeepSeek-V4 Flash is 53.3% cheaper for input tokens ($0.14 vs. $0.30 per 1M tokens) and $0.56 vs. $0.90 for output tokens (2.1x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (0ms faster than Codestral 25.01). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Fable 5

Codestral 25.01 is 85.0% cheaper for input tokens ($0.30 vs. $2.00 per 1M tokens) and $0.90 vs. $8.00 for output tokens (6.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (90ms faster than Fable 5). Fable 5 leads coding benchmarks at 58.0% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Gemini 2.0 Flash

Gemini 2.0 Flash is 66.7% cheaper for input tokens ($0.10 vs. $0.30 per 1M tokens) and $0.40 vs. $0.90 for output tokens (3.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (230ms faster than Gemini 2.0 Flash). Gemini 2.0 Flash leads coding benchmarks at 48.0% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Gemini 3.7 Flash

Gemini 3.7 Flash is 73.3% cheaper for input tokens ($0.08 vs. $0.30 per 1M tokens) and $0.32 vs. $0.90 for output tokens (3.8x cost difference). In terms of operational performance, Gemini 3.7 Flash delivers faster response latency with 75ms TTFT (75ms faster than Codestral 25.01). Gemini 3.7 Flash leads coding benchmarks at 68.2% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs GLM 5.3 Flash

GLM 5.3 Flash is 50.0% cheaper for input tokens ($0.15 vs. $0.30 per 1M tokens) and $0.50 vs. $0.90 for output tokens (2.0x cost difference). In terms of operational performance, GLM 5.3 Flash delivers faster response latency with 130ms TTFT (20ms faster than Codestral 25.01). GLM 5.3 Flash leads coding benchmarks at 54.2% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs OpenAI GPT-4o

Codestral 25.01 is 88.0% cheaper for input tokens ($0.30 vs. $2.50 per 1M tokens) and $0.90 vs. $10.00 for output tokens (8.3x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (130ms faster than OpenAI GPT-4o). Codestral 25.01 leads coding benchmarks at 44.2% SWE-bench vs. 38.8% for OpenAI GPT-4o.

Versus Comparison

Codestral 25.01 vs GPT-5.6 Luna

GPT-5.6 Luna is 40.0% cheaper for input tokens ($0.18 vs. $0.30 per 1M tokens) and $0.72 vs. $0.90 for output tokens (1.7x cost difference). In terms of operational performance, GPT-5.6 Luna delivers faster response latency with 90ms TTFT (60ms faster than Codestral 25.01). GPT-5.6 Luna leads coding benchmarks at 48.5% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs GPT-5.6 Sol

Codestral 25.01 is 96.2% cheaper for input tokens ($0.30 vs. $8.00 per 1M tokens) and $0.90 vs. $32.00 for output tokens (26.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (270ms faster than GPT-5.6 Sol). GPT-5.6 Sol leads coding benchmarks at 79.5% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs GPT-5.6 Terra

Codestral 25.01 is 80.0% cheaper for input tokens ($0.30 vs. $1.50 per 1M tokens) and $0.90 vs. $6.00 for output tokens (5.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (60ms faster than GPT-5.6 Terra). GPT-5.6 Terra leads coding benchmarks at 65.4% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Grok 3

Codestral 25.01 is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $0.90 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (700ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs xAI Grok 4.6

Codestral 25.01 is 85.0% cheaper for input tokens ($0.30 vs. $2.00 per 1M tokens) and $0.90 vs. $6.00 for output tokens (6.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (130ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Llama 3.3 70B Instruct

Llama 3.3 70B Instruct is 40.0% cheaper for input tokens ($0.18 vs. $0.30 per 1M tokens) and $0.40 vs. $0.90 for output tokens (1.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (270ms faster than Llama 3.3 70B Instruct). Codestral 25.01 leads coding benchmarks at 44.2% SWE-bench vs. 38.8% for Llama 3.3 70B Instruct.

Versus Comparison

Codestral 25.01 vs Mistral Large 2

Codestral 25.01 is 85.0% cheaper for input tokens ($0.30 vs. $2.00 per 1M tokens) and $0.90 vs. $6.00 for output tokens (6.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (400ms faster than Mistral Large 2). Codestral 25.01 leads coding benchmarks at 44.2% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Codestral 25.01 vs OpenAI o1

Codestral 25.01 is 98.0% cheaper for input tokens ($0.30 vs. $15.00 per 1M tokens) and $0.90 vs. $60.00 for output tokens (50.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (700ms faster than OpenAI o1). OpenAI o1 leads coding benchmarks at 48.9% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs o3-mini

Codestral 25.01 is 72.7% cheaper for input tokens ($0.30 vs. $1.10 per 1M tokens) and $0.90 vs. $4.40 for output tokens (3.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (1050ms faster than o3-mini). o3-mini leads coding benchmarks at 49.3% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Codestral 25.01 vs Microsoft Phi-4 (14B)

Microsoft Phi-4 (14B) is 60.0% cheaper for input tokens ($0.12 vs. $0.30 per 1M tokens) and $0.36 vs. $0.90 for output tokens (2.5x cost difference). In terms of operational performance, Microsoft Phi-4 (14B) delivers faster response latency with 110ms TTFT (40ms faster than Codestral 25.01). Codestral 25.01 leads coding benchmarks at 44.2% SWE-bench vs. 42.1% for Microsoft Phi-4 (14B).

Versus Comparison

Codestral 25.01 vs Qwen 2.5 72B Instruct

Codestral 25.01 is 14.3% cheaper for input tokens ($0.30 vs. $0.35 per 1M tokens) and $0.90 vs. $0.40 for output tokens (1.2x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (270ms faster than Qwen 2.5 72B Instruct). Codestral 25.01 leads coding benchmarks at 44.2% SWE-bench vs. 44.0% for Qwen 2.5 72B Instruct.

Versus Comparison

Codestral 25.01 vs Qwen 2.5 Max

Qwen 2.5 Max is 6.7% cheaper for input tokens ($0.28 vs. $0.30 per 1M tokens) and $0.84 vs. $0.90 for output tokens (1.1x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (330ms faster than Qwen 2.5 Max). Both models demonstrate comparable coding benchmark scores.

Versus Comparison

Codestral 25.01 vs Qwen 3.8 Flash Next

Qwen 3.8 Flash Next is 60.0% cheaper for input tokens ($0.12 vs. $0.30 per 1M tokens) and $0.48 vs. $0.90 for output tokens (2.5x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 120ms TTFT (30ms faster than Codestral 25.01). Qwen 3.8 Flash Next leads coding benchmarks at 56.8% SWE-bench vs. 44.2% for Codestral 25.01.

Frequently Asked Questions & Query Fan-Out

How much does Codestral 25.01 cost per 1M tokens?

Codestral 25.01 costs $0.30 per million prompt (input) tokens and $0.90 per million completion (output) tokens.

What is the context window for Codestral 25.01?

Codestral 25.01 supports a maximum context window of 262,144 tokens, with a maximum single-generation output of 8,192 tokens.

What are the primary use cases for Codestral 25.01?

256K long-context repository indexing, fill-in-the-middle code completion, and IDE plugin backend.