What are the token costs and operational benchmarks for Claude 3.5 Sonnet?

Claude 3.5 Sonnet is priced at $3.00 per million input tokens and $15.00 per million output tokens. It features a 204.8k token context window, an average response latency of 320ms TTFT, and achieves 63.7% on SWE-bench Verified and 75.9% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $3.00
Output / 1M $15.00
Context Limit 204.8k
TTFT Latency 320ms
SWE-bench 63.7%
Throughput 72 t/s

Architectural Overview

Frontier foundation model developed by Anthropic featuring Dense Multimodal Transformer architecture.

Optimal Production Use Cases

Full-stack code generation, production AI agents, complex visual diagram parsing, and enterprise document workflows.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

Claude 3.5 Sonnet $5.45

$3/1M in · $15/1M out

GPT-5.6 Luna $67.50

$0.18/1M in · $0.72/1M out

Projected Monthly Cost Reduction
$62.05 / mo
(91.9% lower cost)

Head-to-Head Comparisons Involving Claude 3.5 Sonnet

Versus Comparison

Claude 3.5 Sonnet vs Claude 3.5 Haiku

Claude 3.5 Haiku is 73.3% cheaper for input tokens ($0.80 vs. $3.00 per 1M tokens) and $4.00 vs. $15.00 for output tokens (3.8x cost difference). In terms of operational performance, Claude 3.5 Haiku delivers faster response latency with 140ms TTFT (180ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 40.6% for Claude 3.5 Haiku.

Versus Comparison

Claude 3.5 Sonnet vs Claude 3.7 Sonnet

Both models share identical input pricing at $3.00 per 1M tokens. In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (330ms faster than Claude 3.7 Sonnet). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 63.7% for Claude 3.5 Sonnet.

Versus Comparison

Claude 3.5 Sonnet vs Claude Opus 5

Claude 3.5 Sonnet is 40.0% cheaper for input tokens ($3.00 vs. $5.00 per 1M tokens) and $15.00 vs. $25.00 for output tokens (1.7x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (20ms faster than Claude Opus 5). Claude Opus 5 leads coding benchmarks at 82.4% SWE-bench vs. 63.7% for Claude 3.5 Sonnet.

Versus Comparison

Claude 3.5 Sonnet vs Codestral 25.01

Codestral 25.01 is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $0.90 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (170ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Claude 3.5 Sonnet vs Composer 2.5

Composer 2.5 is 40.0% cheaper for input tokens ($1.80 vs. $3.00 per 1M tokens) and $7.20 vs. $15.00 for output tokens (1.7x cost difference). In terms of operational performance, Composer 2.5 delivers faster response latency with 180ms TTFT (140ms faster than Claude 3.5 Sonnet). Composer 2.5 leads coding benchmarks at 74.6% SWE-bench vs. 63.7% for Claude 3.5 Sonnet.

Versus Comparison

Claude 3.5 Sonnet vs DeepSeek-R1

DeepSeek-R1 is 81.7% cheaper for input tokens ($0.55 vs. $3.00 per 1M tokens) and $2.19 vs. $15.00 for output tokens (5.5x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (1480ms faster than DeepSeek-R1). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

Claude 3.5 Sonnet vs DeepSeek-V3

DeepSeek-V3 is 95.3% cheaper for input tokens ($0.14 vs. $3.00 per 1M tokens) and $0.28 vs. $15.00 for output tokens (21.4x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (20ms faster than DeepSeek-V3). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 42.0% for DeepSeek-V3.

Versus Comparison

Claude 3.5 Sonnet vs DeepSeek-V4 Flash

DeepSeek-V4 Flash is 95.3% cheaper for input tokens ($0.14 vs. $3.00 per 1M tokens) and $0.56 vs. $15.00 for output tokens (21.4x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (170ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.

Versus Comparison

Claude 3.5 Sonnet vs Fable 5

Fable 5 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $8.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, Fable 5 delivers faster response latency with 240ms TTFT (80ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 58.0% for Fable 5.

Versus Comparison

Claude 3.5 Sonnet vs Gemini 2.0 Flash

Gemini 2.0 Flash is 96.7% cheaper for input tokens ($0.10 vs. $3.00 per 1M tokens) and $0.40 vs. $15.00 for output tokens (30.0x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (60ms faster than Gemini 2.0 Flash). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 48.0% for Gemini 2.0 Flash.

Versus Comparison

Claude 3.5 Sonnet vs Gemini 3.7 Flash

Gemini 3.7 Flash is 97.3% cheaper for input tokens ($0.08 vs. $3.00 per 1M tokens) and $0.32 vs. $15.00 for output tokens (37.5x cost difference). In terms of operational performance, Gemini 3.7 Flash delivers faster response latency with 75ms TTFT (245ms faster than Claude 3.5 Sonnet). Gemini 3.7 Flash leads coding benchmarks at 68.2% SWE-bench vs. 63.7% for Claude 3.5 Sonnet.

Versus Comparison

Claude 3.5 Sonnet vs GLM 5.3 Flash

GLM 5.3 Flash is 95.0% cheaper for input tokens ($0.15 vs. $3.00 per 1M tokens) and $0.50 vs. $15.00 for output tokens (20.0x cost difference). In terms of operational performance, GLM 5.3 Flash delivers faster response latency with 130ms TTFT (190ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 54.2% for GLM 5.3 Flash.

Versus Comparison

Claude 3.5 Sonnet vs OpenAI GPT-4o

OpenAI GPT-4o is 16.7% cheaper for input tokens ($2.50 vs. $3.00 per 1M tokens) and $10.00 vs. $15.00 for output tokens (1.2x cost difference). In terms of operational performance, OpenAI GPT-4o delivers faster response latency with 280ms TTFT (40ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 38.8% for OpenAI GPT-4o.

Versus Comparison

Claude 3.5 Sonnet vs GPT-5.6 Luna

GPT-5.6 Luna is 94.0% cheaper for input tokens ($0.18 vs. $3.00 per 1M tokens) and $0.72 vs. $15.00 for output tokens (16.7x cost difference). In terms of operational performance, GPT-5.6 Luna delivers faster response latency with 90ms TTFT (230ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 48.5% for GPT-5.6 Luna.

Versus Comparison

Claude 3.5 Sonnet vs GPT-5.6 Sol

Claude 3.5 Sonnet is 62.5% cheaper for input tokens ($3.00 vs. $8.00 per 1M tokens) and $15.00 vs. $32.00 for output tokens (2.7x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (100ms faster than GPT-5.6 Sol). GPT-5.6 Sol leads coding benchmarks at 79.5% SWE-bench vs. 63.7% for Claude 3.5 Sonnet.

Versus Comparison

Claude 3.5 Sonnet vs GPT-5.6 Terra

GPT-5.6 Terra is 50.0% cheaper for input tokens ($1.50 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (2.0x cost difference). In terms of operational performance, GPT-5.6 Terra delivers faster response latency with 210ms TTFT (110ms faster than Claude 3.5 Sonnet). GPT-5.6 Terra leads coding benchmarks at 65.4% SWE-bench vs. 63.7% for Claude 3.5 Sonnet.

Versus Comparison

Claude 3.5 Sonnet vs Grok 3

Both models share identical input pricing at $3.00 per 1M tokens. In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (530ms faster than Grok 3). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Claude 3.5 Sonnet vs xAI Grok 4.6

xAI Grok 4.6 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (40ms faster than Claude 3.5 Sonnet). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 63.7% for Claude 3.5 Sonnet.

Versus Comparison

Claude 3.5 Sonnet vs Llama 3.3 70B Instruct

Llama 3.3 70B Instruct is 94.0% cheaper for input tokens ($0.18 vs. $3.00 per 1M tokens) and $0.40 vs. $15.00 for output tokens (16.7x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (100ms faster than Llama 3.3 70B Instruct). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 38.8% for Llama 3.3 70B Instruct.

Versus Comparison

Claude 3.5 Sonnet vs Mistral Large 2

Mistral Large 2 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (230ms faster than Mistral Large 2). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Claude 3.5 Sonnet vs OpenAI o1

Claude 3.5 Sonnet is 80.0% cheaper for input tokens ($3.00 vs. $15.00 per 1M tokens) and $15.00 vs. $60.00 for output tokens (5.0x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (530ms faster than OpenAI o1). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 48.9% for OpenAI o1.

Versus Comparison

Claude 3.5 Sonnet vs o3-mini

o3-mini is 63.3% cheaper for input tokens ($1.10 vs. $3.00 per 1M tokens) and $4.40 vs. $15.00 for output tokens (2.7x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (880ms faster than o3-mini). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 49.3% for o3-mini.

Versus Comparison

Claude 3.5 Sonnet vs Microsoft Phi-4 (14B)

Microsoft Phi-4 (14B) is 96.0% cheaper for input tokens ($0.12 vs. $3.00 per 1M tokens) and $0.36 vs. $15.00 for output tokens (25.0x cost difference). In terms of operational performance, Microsoft Phi-4 (14B) delivers faster response latency with 110ms TTFT (210ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 42.1% for Microsoft Phi-4 (14B).

Versus Comparison

Claude 3.5 Sonnet vs Qwen 2.5 72B Instruct

Qwen 2.5 72B Instruct is 88.3% cheaper for input tokens ($0.35 vs. $3.00 per 1M tokens) and $0.40 vs. $15.00 for output tokens (8.6x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (100ms faster than Qwen 2.5 72B Instruct). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 44.0% for Qwen 2.5 72B Instruct.

Versus Comparison

Claude 3.5 Sonnet vs Qwen 2.5 Max

Qwen 2.5 Max is 90.7% cheaper for input tokens ($0.28 vs. $3.00 per 1M tokens) and $0.84 vs. $15.00 for output tokens (10.7x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (160ms faster than Qwen 2.5 Max). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 44.2% for Qwen 2.5 Max.

Versus Comparison

Claude 3.5 Sonnet vs Qwen 3.8 Flash Next

Qwen 3.8 Flash Next is 96.0% cheaper for input tokens ($0.12 vs. $3.00 per 1M tokens) and $0.48 vs. $15.00 for output tokens (25.0x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 120ms TTFT (200ms faster than Claude 3.5 Sonnet). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 56.8% for Qwen 3.8 Flash Next.

Frequently Asked Questions & Query Fan-Out

How much does Claude 3.5 Sonnet cost per 1M tokens?

Claude 3.5 Sonnet costs $3.00 per million prompt (input) tokens and $15.00 per million completion (output) tokens.

What is the context window for Claude 3.5 Sonnet?

Claude 3.5 Sonnet supports a maximum context window of 204,800 tokens, with a maximum single-generation output of 8,192 tokens.

What are the primary use cases for Claude 3.5 Sonnet?

Full-stack code generation, production AI agents, complex visual diagram parsing, and enterprise document workflows.