What are the token costs and operational benchmarks for Grok 3?

Grok 3 is priced at $3.00 per million input tokens and $15.00 per million output tokens. It features a 1,024k token context window, an average response latency of 850ms TTFT, and achieves 58.5% on SWE-bench Verified and 82.3% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $3.00
Output / 1M $15.00
Context Limit 1,024k
TTFT Latency 850ms
SWE-bench 58.5%
Throughput 70 t/s

Architectural Overview

Frontier foundation model developed by xAI featuring Dense Transformer architecture.

Optimal Production Use Cases

Frontier scientific reasoning, real-time web exploration, and advanced mathematical problem solving.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

Grok 3 $5.45

$3/1M in · $15/1M out

GPT-5.6 Luna $67.50

$0.18/1M in · $0.72/1M out

Projected Monthly Cost Reduction
$62.05 / mo
(91.9% lower cost)

Head-to-Head Comparisons Involving Grok 3

Versus Comparison

Grok 3 vs Claude 3.5 Haiku

Claude 3.5 Haiku is 73.3% cheaper for input tokens ($0.80 vs. $3.00 per 1M tokens) and $4.00 vs. $15.00 for output tokens (3.8x cost difference). In terms of operational performance, Claude 3.5 Haiku delivers faster response latency with 140ms TTFT (710ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 40.6% for Claude 3.5 Haiku.

Versus Comparison

Grok 3 vs Claude 3.5 Sonnet

Both models share identical input pricing at $3.00 per 1M tokens. In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (530ms faster than Grok 3). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs Claude 3.7 Sonnet

Both models share identical input pricing at $3.00 per 1M tokens. In terms of operational performance, Claude 3.7 Sonnet delivers faster response latency with 650ms TTFT (200ms faster than Grok 3). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs Claude Opus 5

Grok 3 is 40.0% cheaper for input tokens ($3.00 vs. $5.00 per 1M tokens) and $15.00 vs. $25.00 for output tokens (1.7x cost difference). In terms of operational performance, Claude Opus 5 delivers faster response latency with 340ms TTFT (510ms faster than Grok 3). Claude Opus 5 leads coding benchmarks at 82.4% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs Codestral 25.01

Codestral 25.01 is 90.0% cheaper for input tokens ($0.30 vs. $3.00 per 1M tokens) and $0.90 vs. $15.00 for output tokens (10.0x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (700ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

Grok 3 vs Composer 2.5

Composer 2.5 is 40.0% cheaper for input tokens ($1.80 vs. $3.00 per 1M tokens) and $7.20 vs. $15.00 for output tokens (1.7x cost difference). In terms of operational performance, Composer 2.5 delivers faster response latency with 180ms TTFT (670ms faster than Grok 3). Composer 2.5 leads coding benchmarks at 74.6% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs DeepSeek-R1

DeepSeek-R1 is 81.7% cheaper for input tokens ($0.55 vs. $3.00 per 1M tokens) and $2.19 vs. $15.00 for output tokens (5.5x cost difference). In terms of operational performance, Grok 3 delivers faster response latency with 850ms TTFT (950ms faster than DeepSeek-R1). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

Grok 3 vs DeepSeek-V3

DeepSeek-V3 is 95.3% cheaper for input tokens ($0.14 vs. $3.00 per 1M tokens) and $0.28 vs. $15.00 for output tokens (21.4x cost difference). In terms of operational performance, DeepSeek-V3 delivers faster response latency with 340ms TTFT (510ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 42.0% for DeepSeek-V3.

Versus Comparison

Grok 3 vs DeepSeek-V4 Flash

DeepSeek-V4 Flash is 95.3% cheaper for input tokens ($0.14 vs. $3.00 per 1M tokens) and $0.56 vs. $15.00 for output tokens (21.4x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (700ms faster than Grok 3). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs Fable 5

Fable 5 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $8.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, Fable 5 delivers faster response latency with 240ms TTFT (610ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 58.0% for Fable 5.

Versus Comparison

Grok 3 vs Gemini 2.0 Flash

Gemini 2.0 Flash is 96.7% cheaper for input tokens ($0.10 vs. $3.00 per 1M tokens) and $0.40 vs. $15.00 for output tokens (30.0x cost difference). In terms of operational performance, Gemini 2.0 Flash delivers faster response latency with 380ms TTFT (470ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 48.0% for Gemini 2.0 Flash.

Versus Comparison

Grok 3 vs Gemini 3.7 Flash

Gemini 3.7 Flash is 97.3% cheaper for input tokens ($0.08 vs. $3.00 per 1M tokens) and $0.32 vs. $15.00 for output tokens (37.5x cost difference). In terms of operational performance, Gemini 3.7 Flash delivers faster response latency with 75ms TTFT (775ms faster than Grok 3). Gemini 3.7 Flash leads coding benchmarks at 68.2% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs GLM 5.3 Flash

GLM 5.3 Flash is 95.0% cheaper for input tokens ($0.15 vs. $3.00 per 1M tokens) and $0.50 vs. $15.00 for output tokens (20.0x cost difference). In terms of operational performance, GLM 5.3 Flash delivers faster response latency with 130ms TTFT (720ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 54.2% for GLM 5.3 Flash.

Versus Comparison

Grok 3 vs OpenAI GPT-4o

OpenAI GPT-4o is 16.7% cheaper for input tokens ($2.50 vs. $3.00 per 1M tokens) and $10.00 vs. $15.00 for output tokens (1.2x cost difference). In terms of operational performance, OpenAI GPT-4o delivers faster response latency with 280ms TTFT (570ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 38.8% for OpenAI GPT-4o.

Versus Comparison

Grok 3 vs GPT-5.6 Luna

GPT-5.6 Luna is 94.0% cheaper for input tokens ($0.18 vs. $3.00 per 1M tokens) and $0.72 vs. $15.00 for output tokens (16.7x cost difference). In terms of operational performance, GPT-5.6 Luna delivers faster response latency with 90ms TTFT (760ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 48.5% for GPT-5.6 Luna.

Versus Comparison

Grok 3 vs GPT-5.6 Sol

Grok 3 is 62.5% cheaper for input tokens ($3.00 vs. $8.00 per 1M tokens) and $15.00 vs. $32.00 for output tokens (2.7x cost difference). In terms of operational performance, GPT-5.6 Sol delivers faster response latency with 420ms TTFT (430ms faster than Grok 3). GPT-5.6 Sol leads coding benchmarks at 79.5% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs GPT-5.6 Terra

GPT-5.6 Terra is 50.0% cheaper for input tokens ($1.50 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (2.0x cost difference). In terms of operational performance, GPT-5.6 Terra delivers faster response latency with 210ms TTFT (640ms faster than Grok 3). GPT-5.6 Terra leads coding benchmarks at 65.4% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs xAI Grok 4.6

xAI Grok 4.6 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (570ms faster than Grok 3). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

Grok 3 vs Llama 3.3 70B Instruct

Llama 3.3 70B Instruct is 94.0% cheaper for input tokens ($0.18 vs. $3.00 per 1M tokens) and $0.40 vs. $15.00 for output tokens (16.7x cost difference). In terms of operational performance, Llama 3.3 70B Instruct delivers faster response latency with 420ms TTFT (430ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 38.8% for Llama 3.3 70B Instruct.

Versus Comparison

Grok 3 vs Mistral Large 2

Mistral Large 2 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, Mistral Large 2 delivers faster response latency with 550ms TTFT (300ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Grok 3 vs OpenAI o1

Grok 3 is 80.0% cheaper for input tokens ($3.00 vs. $15.00 per 1M tokens) and $15.00 vs. $60.00 for output tokens (5.0x cost difference). In terms of operational performance, OpenAI o1 delivers faster response latency with 850ms TTFT (0ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 48.9% for OpenAI o1.

Versus Comparison

Grok 3 vs o3-mini

o3-mini is 63.3% cheaper for input tokens ($1.10 vs. $3.00 per 1M tokens) and $4.40 vs. $15.00 for output tokens (2.7x cost difference). In terms of operational performance, Grok 3 delivers faster response latency with 850ms TTFT (350ms faster than o3-mini). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 49.3% for o3-mini.

Versus Comparison

Grok 3 vs Microsoft Phi-4 (14B)

Microsoft Phi-4 (14B) is 96.0% cheaper for input tokens ($0.12 vs. $3.00 per 1M tokens) and $0.36 vs. $15.00 for output tokens (25.0x cost difference). In terms of operational performance, Microsoft Phi-4 (14B) delivers faster response latency with 110ms TTFT (740ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 42.1% for Microsoft Phi-4 (14B).

Versus Comparison

Grok 3 vs Qwen 2.5 72B Instruct

Qwen 2.5 72B Instruct is 88.3% cheaper for input tokens ($0.35 vs. $3.00 per 1M tokens) and $0.40 vs. $15.00 for output tokens (8.6x cost difference). In terms of operational performance, Qwen 2.5 72B Instruct delivers faster response latency with 420ms TTFT (430ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 44.0% for Qwen 2.5 72B Instruct.

Versus Comparison

Grok 3 vs Qwen 2.5 Max

Qwen 2.5 Max is 90.7% cheaper for input tokens ($0.28 vs. $3.00 per 1M tokens) and $0.84 vs. $15.00 for output tokens (10.7x cost difference). In terms of operational performance, Qwen 2.5 Max delivers faster response latency with 480ms TTFT (370ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 44.2% for Qwen 2.5 Max.

Versus Comparison

Grok 3 vs Qwen 3.8 Flash Next

Qwen 3.8 Flash Next is 96.0% cheaper for input tokens ($0.12 vs. $3.00 per 1M tokens) and $0.48 vs. $15.00 for output tokens (25.0x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 120ms TTFT (730ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 56.8% for Qwen 3.8 Flash Next.

Frequently Asked Questions & Query Fan-Out

How much does Grok 3 cost per 1M tokens?

Grok 3 costs $3.00 per million prompt (input) tokens and $15.00 per million completion (output) tokens.

What is the context window for Grok 3?

Grok 3 supports a maximum context window of 1,024,000 tokens, with a maximum single-generation output of 32,768 tokens.

What are the primary use cases for Grok 3?

Frontier scientific reasoning, real-time web exploration, and advanced mathematical problem solving.