Qwen 3.8 27B
Apache 2.0 Open SourceDeveloped by Alibaba Cloud · Released 2026-02-15
What are the token costs and operational benchmarks for Qwen 3.8 27B?
Architectural Overview
Frontier foundation model developed by Alibaba Cloud featuring Dense Transformer architecture.
Optimal Production Use Cases
High-performance 27B open weights optimized for local deployment, code generation, and low-latency production pipelines.
Interactive Monthly Token Economics & ROI Forecaster
Model your expected production workload across prompt (input) and completion (output) tokens with prompt caching economics.
Estimated Monthly Spend
$0.20/1M in · $0.60/1M out
Prompt Cache: $0.050/1M read · $0.20/1M write
$0.10/1M in · $0.35/1M out
Prompt Cache: $0.020/1M read · $0.10/1M write
Head-to-Head Comparisons Involving Qwen 3.8 27B
Qwen 3.8 27B vs Claude 3.5 Haiku
Qwen 3.8 27B is 75.0% cheaper for input tokens ($0.20 vs. $0.80 per 1M tokens) and $0.60 vs. $4.00 for output tokens (4.0x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (70ms faster than Claude 3.5 Haiku). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 40.6% for Claude 3.5 Haiku.
Qwen 3.8 27B vs Claude 3.5 Sonnet
Qwen 3.8 27B is 93.3% cheaper for input tokens ($0.20 vs. $3.00 per 1M tokens) and $0.60 vs. $15.00 for output tokens (15.0x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (190ms faster than Claude 3.5 Sonnet). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 49.0% for Claude 3.5 Sonnet.
Qwen 3.8 27B vs Claude 3.7 Sonnet
Qwen 3.8 27B is 93.3% cheaper for input tokens ($0.20 vs. $3.00 per 1M tokens) and $0.60 vs. $15.00 for output tokens (15.0x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (530ms faster than Claude 3.7 Sonnet). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 58.4% for Qwen 3.8 27B.
Qwen 3.8 27B vs DeepSeek R1
Qwen 3.8 27B is 63.6% cheaper for input tokens ($0.20 vs. $0.55 per 1M tokens) and $0.60 vs. $2.19 for output tokens (2.8x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (1680ms faster than DeepSeek-R1). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 49.2% for DeepSeek-R1.
Qwen 3.8 27B vs DeepSeek V3
**DeepSeek V3 provides superior cost efficiency, while Qwen 3.8 27B dominates coding capability and latency.** DeepSeek V3 is 30.0% cheaper for inputs ($0.14 vs. $0.20 per 1M) and outputs ($0.28 vs. $0.60). Conversely, Qwen 3.8 27B achieves a superior 58.4% SWE-bench score over 42.0%, delivering faster response times with a 120ms TTFT (290ms faster).
Qwen 3.8 27B vs DeepSeek-V4 Flash Vision Exp
DeepSeek-V4 Flash Vision Exp is 30.0% cheaper for input tokens ($0.14 vs. $0.20 per 1M tokens) and $0.56 vs. $0.60 for output tokens (1.4x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (20ms faster than DeepSeek-V4 Flash Vision Exp). DeepSeek-V4 Flash Vision Exp leads coding benchmarks at 64.2% SWE-bench vs. 58.4% for Qwen 3.8 27B.
Qwen 3.8 27B vs Gemini 2.0 Flash
Gemini 2.0 Flash is 50.0% cheaper for input tokens ($0.10 vs. $0.20 per 1M tokens) and $0.40 vs. $0.60 for output tokens (2.0x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (100ms faster than Gemini 2.0 Flash). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 28.0% for Gemini 2.0 Flash.
Qwen 3.8 27B vs Gemini 2.5 Flash
Qwen 3.8 27B is 33.3% cheaper for input tokens ($0.20 vs. $0.30 per 1M tokens) and $0.60 vs. $2.50 for output tokens (1.5x cost difference). In terms of operational performance, Gemini 2.5 Flash delivers faster response latency with 110ms TTFT (10ms faster than Qwen 3.8 27B). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 58.0% for Gemini 2.5 Flash.
Qwen 3.8 27B vs Gemini 2.5 Pro
Qwen 3.8 27B is 84.0% cheaper for input tokens ($0.20 vs. $1.25 per 1M tokens) and $0.60 vs. $10.00 for output tokens (6.2x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (230ms faster than Gemini 2.5 Pro). Gemini 2.5 Pro leads coding benchmarks at 65.0% SWE-bench vs. 58.4% for Qwen 3.8 27B.
Qwen 3.8 27B vs GLM-5.3-Flash
GLM-5.3-Flash is 40.0% cheaper for input tokens ($0.12 vs. $0.20 per 1M tokens) and $0.40 vs. $0.60 for output tokens (1.7x cost difference). In terms of operational performance, GLM-5.3-Flash delivers faster response latency with 90ms TTFT (30ms faster than Qwen 3.8 27B). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 52.8% for GLM-5.3-Flash.
Qwen 3.8 27B vs GLM-5.3
Qwen 3.8 27B is 42.9% cheaper for input tokens ($0.20 vs. $0.35 per 1M tokens) and $0.60 vs. $1.00 for output tokens (1.7x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (40ms faster than GLM-5.3). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 56.0% for GLM-5.3.
Qwen 3.8 27B vs GPT-4.5
**Qwen 3.8 27B delivers vastly superior value, slashing costs by 99.7% while outperforming GPT-4.5 in key capabilities.** At $0.20 input and $0.60 output per 1M tokens, Qwen is up to 375x cheaper than GPT-4.5 ($75.00/$150.00). It also responds 730ms faster with a 120ms TTFT and dominates SWE-bench coding at 58.4% versus 38.0%.
Qwen 3.8 27B vs GPT-4o Mini
GPT-4o Mini is 25.0% cheaper for input tokens ($0.15 vs. $0.20 per 1M tokens) and $0.60 vs. $0.60 for output tokens (1.3x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (100ms faster than GPT-4o Mini). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 41.0% for GPT-4o Mini.
Qwen 3.8 27B vs GPT-4o
Qwen 3.8 27B is 92.0% cheaper for input tokens ($0.20 vs. $2.50 per 1M tokens) and $0.60 vs. $10.00 for output tokens (12.5x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (160ms faster than GPT-4o). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 48.0% for GPT-4o.
Qwen 3.8 27B vs Grok 3 Mini
Qwen 3.8 27B is 33.3% cheaper for input tokens ($0.20 vs. $0.30 per 1M tokens) and $0.60 vs. $1.20 for output tokens (1.5x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (30ms faster than Grok 3 Mini). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 54.0% for Grok 3 Mini.
Qwen 3.8 27B vs Grok 3
**Qwen 3.8 27B offers overwhelming cost and speed advantages, whereas Grok 3 dominates top-tier coding performance.** Qwen is 93.3% cheaper for input tokens ($0.20 vs. $3.00 per 1M) and outputs ($0.60 vs. $15.00), clocking a 120ms TTFT (500ms faster than Grok 3). However, Grok 3 leads SWE-bench at 65.8% versus Qwen’s 58.4%.
Qwen 3.8 27B vs Llama 3.3 70B Instruct
**Llama 3.3 70B Instruct offers superior cost efficiency, while Qwen 3.8 27B dominates in speed and coding capability.** Llama cuts pricing by 40.0% on inputs ($0.12 vs. $0.20/1M) and outputs ($0.30 vs. $0.60/1M). Conversely, Qwen delivers a faster 120ms TTFT (330ms advantage) and higher coding performance, scoring 58.4% versus 42.5% on SWE-bench.
Qwen 3.8 27B vs Mistral Large 2
**Qwen 3.8 27B delivers superior capability at a 90.0% lower cost than Mistral Large 2.** Qwen costs just $0.20 input and $0.60 output per 1M tokens versus Mistral's $2.00 and $6.00. Furthermore, Qwen achieves a faster 120ms TTFT (a 400ms advantage) and leads coding with a 58.4% SWE-bench score over Mistral’s 40.0%.
Qwen 3.8 27B vs Mistral Large 2 (2411)
Qwen 3.8 27B is 90.0% cheaper for input tokens ($0.20 vs. $2.00 per 1M tokens) and $0.60 vs. $6.00 for output tokens (10.0x cost difference). In terms of operational performance, Qwen 3.8 27B delivers faster response latency with 120ms TTFT (330ms faster than Mistral Large 2411). Qwen 3.8 27B leads coding benchmarks at 58.4% SWE-bench vs. 35.7% for Mistral Large 2411.
Qwen 3.8 27B vs o1
**Qwen 3.8 27B delivers superior developer performance at an overwhelmingly lower price than o1.** Qwen is 98.7% cheaper on input tokens ($0.20 vs. $15.00/1M) and output tokens ($0.60 vs. $60.00/1M). It drastically improves response speed with a 120ms TTFT (2480ms faster) and leads coding benchmarks with 58.4% on SWE-bench versus 48.9% for o1.
Qwen 3.8 27B vs OpenAI o3-mini
**Qwen 3.8 27B decisively outperforms OpenAI o3-mini in cost efficiency, responsiveness, and coding capability.** Qwen is 81.8% cheaper for inputs at $0.20 versus $1.10 per 1M tokens, and $0.60 versus $4.40 for outputs. It delivers a 120ms TTFT (1080ms faster) and leads SWE-bench with 58.4% against o3-mini’s 49.3%.
Qwen 3.8 27B vs Qwen 3.8 Flash Next
Qwen 3.8 Flash Next is 50.0% cheaper for input tokens ($0.10 vs. $0.20 per 1M tokens) and $0.35 vs. $0.60 for output tokens (2.0x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 95ms TTFT (25ms faster than Qwen 3.8 27B). Qwen 3.8 Flash Next leads coding benchmarks at 61.0% SWE-bench vs. 58.4% for Qwen 3.8 27B.
Frequently Asked Questions & Query Fan-Out
How much does Qwen 3.8 27B cost per 1M tokens?
Qwen 3.8 27B costs $0.20 per million prompt (input) tokens and $0.60 per million completion (output) tokens.
What is the context window for Qwen 3.8 27B?
Qwen 3.8 27B supports a maximum context window of 131,072 tokens, with a maximum single-generation output of 16,384 tokens.
What are the primary use cases for Qwen 3.8 27B?
High-performance 27B open weights optimized for local deployment, code generation, and low-latency production pipelines.