// benchmark · coding-llms

The LLM coding benchmark

Every coding model that mattered, from the dawn of the agentic-coding era (late 2024) to today: when it shipped, how big a window, how well it closes real GitHub issues, and what it costs. Click any column to sort — try Released for the full timeline.

open weights you can self-hostclosed API onlySWE-bench Verified = real-issue solve rate
ModelWeightsReleasedContextSWE-bench VerifiedSWE-bench ProInput /MOutput /MBest for
Claude Opus 5
Anthropic
closed 2026-07 1M n/a n/a $5 $25 Newest Anthropic flagship; 1M context; no published SWE-bench Verified run yet
Claude Opus 5 (fast mode)
Anthropic
closed 2026-07 1M n/a n/a $10 $50 Low-latency Opus 5 tier, 2x standard price; no published Verified run yet
Claude Mythos 5
Anthropic
closed 2026-06 1M n/a n/a $10 $50 Safeguards-lifted Mythos tier; restricted rollout
Claude Fable 5
Anthropic
closed 2026-06 1M 95.0% 80.3% $10 $50 Mythos-class frontier (restored after US ban)
GPT-5.6 Sol
OpenAI
closed 2026-06 400K 96.2% 64.6% $5 $30 Flagship; Terminal-Bench leader; Verified score from vals.ai independent run (OpenAI retired the benchmark)
GPT-5.6 Terra
OpenAI
closed 2026-06 400K n/a 63.4% $2.5 $15 Lower-cost, ~GPT-5.5 class; no published Verified run
GPT-5.6 Luna
OpenAI
closed 2026-06 400K 93.0% 62.7% $1 $6 Fastest, cheapest GPT-5.6 tier; Verified score from vals.ai independent run
Grok 4.5
xAI
closed 2026-07 500K n/a 64.7% $2 $6 Newest xAI flagship; SWE-bench Pro 64.7%, no Verified score
Gemini 3.6 Flash
Google
closed 2026-07 1M n/a 58.7% $1.5 $7.5 Agent-tuned Flash, ~17% fewer output tokens; SWE-bench Pro 58.7%, no Verified score
Claude Sonnet 5
Anthropic
closed 2026-06 1M 72.7% n/a $3 $15 Strong agentic build follow-through, pricier per task than 4.6
Claude Opus 4.8
Anthropic
closed 2026-05 1M 78.0% 69.2% $5 $25 Frontier agentic coding
Claude Opus 4.8 (fast mode)
Anthropic
closed 2026-05 1M 78.0% n/a $10 $50 Preview low-latency tier, 2x standard Opus 4.8 price
GLM-5.2
Zhipu AI
open 2026-06 200K 78.0% n/a $1 $3 Latest open frontier (indicative)
GLM-5
Zhipu AI
open 2026-04 200K 77.8% n/a $1 $3 Strong open coder (Z.ai)
GPT-5.5 Codex
OpenAI
closed 2026-04 400K 77.0% n/a $5 $30 Tuned for the Codex agent
Claude Opus 4.7
Anthropic
closed 2026-03 1M 77.0% n/a $5 $25 Previous Opus flagship
Claude Opus 4.6
Anthropic
closed 2025-12 1M 77.0% n/a $5 $25 Adaptive-thinking Opus
GPT-5.5
OpenAI
closed 2026-04 400K 76.0% 58.6% $5 $30 OpenAI flagship
Claude Opus 4.5
Anthropic
closed 2025-11 200K 76.0% n/a $5 $25 First effort-dial Opus
Claude Opus 4.1
Anthropic
closed 2025-08 200K 74.0% n/a $15 $75 Older frontier Opus
GPT-5.4
OpenAI
closed 2026-03 400K 74.0% n/a $2.5 $15 Cheaper GPT-5 tier
Qwen3.5 Plus
Alibaba
closed 2026-02 1M 76.4% n/a $0.4 $2.4 Cheap 1M-context hosted coder
Grok 4.20
xAI
closed 2026-03 2M 76.7% n/a $2 $6 2M-context Grok, superseded by 4.3/4.5
GPT-5.4 Mini
OpenAI
closed 2026-03 400K 48.2% n/a $0.75 $4.5 Cheap GPT-5.4 tier, 128k output
Claude Sonnet 4.6
Anthropic
closed 2026-01 1M 73.0% n/a $3 $15 Best price/perf for agents
GPT-5 Codex
OpenAI
closed 2025-09 400K 73.0% n/a $1.25 $10 First Codex-tuned GPT-5
Qwen3.7-Max
Alibaba
closed 2026-03 1M 72.5% n/a $2.5 $7.5 1M-context agent model
Claude Sonnet 4.5
Anthropic
closed 2025-09 200K 72.0% n/a $3 $15 Strong cheaper agent model
Claude Opus 4
Anthropic
closed 2025-05 200K 72.0% n/a $15 $75 First Claude 4 Opus
Claude Sonnet 4
Anthropic
closed 2025-05 200K 71.0% n/a $3 $15 First Claude 4 Sonnet
GPT-5
OpenAI
closed 2025-08 400K 72.0% n/a $1.25 $10 The GPT-5 reasoning leap
Gemini 3.1 Pro
Google
closed 2026-02 1M 72.0% n/a $2 $12 1M context, tiered pricing
Grok 4.3
xAI
closed 2026-05 256K 71.0% n/a $1.25 $2.5 Cheap frontier, fast
Grok Build 0.1
xAI
closed 2026-05 256K 70.8% n/a $1 $2 Coding-agent-tuned Grok, preview
Gemini 3 Pro
Google
closed 2025-11 1M 71.0% n/a $2 $12 The Gemini 3 jump
o3
OpenAI
closed 2025-04 200K 69.0% n/a $2 $8 The o-series reasoning peak
Kimi K2.6
Moonshot AI
open 2026-04 256K 68.0% n/a $0.95 $4 Open agentic coder, 1T MoE
DeepSeek V4 Pro
DeepSeek
open 2026-04 1M 80.6% n/a $0.44 $0.87 Highest open-weight SWE-bench Verified score, very cheap
o4-mini
OpenAI
closed 2025-04 200K 68.0% n/a $1.1 $4.4 Cheap strong reasoner
Grok 4
xAI
closed 2025-07 256K 68.0% n/a $3 $15 The frontier Grok
GLM-4.6
Zhipu AI
open 2025-10 200K 68.0% n/a $0.6 $2.2 Stronger open GLM coder
Qwen3 Max
Alibaba
closed 2025-09 256K 67.0% n/a $0.78 $3.9 First Qwen3 Max
Qwen3-Coder 480B
Alibaba
open 2025-07 256K 66.0% n/a self-host self-host Top open code specialist
DeepSeek V3.1
DeepSeek
open 2025-08 128K 66.0% n/a $0.27 $1.1 Open frontier (V3 line)
Kimi K2
Moonshot AI
open 2025-07 256K 65.8% n/a $0.55 $2.2 The original K2
Gemini 2.5 Pro
Google
closed 2025-03 1M 64.0% n/a $1.25 $10 The long-context breakout
GLM-4.5
Zhipu AI
open 2025-07 128K 64.0% n/a $0.6 $2.2 The GLM open-coder breakout
MiniMax M3
MiniMax
open 2026-06 200K 63.0% n/a $0.3 $1.2 Cheap, agentic, open
Claude Sonnet 3.7
Anthropic
closed 2025-02 200K 62.0% n/a $3 $15 The agentic-coding breakout
DeepSeek V4 Flash
DeepSeek
open 2026-04 1M 60.0% n/a $0.14 $0.28 Cheapest frontier-class
MiniMax M2
MiniMax
open 2025-10 200K 60.0% n/a $0.3 $1.2 Earlier M-series open MoE
Mistral Medium 3.5
Mistral
closed 2026-04 262K 77.6% n/a $1.5 $7.5 Mistral's strongest current model
Mistral Large 3
Mistral
closed 2026-04 256K 58.0% n/a $0.5 $1.5 Cheap European flagship
Mistral Small 26.03
Mistral
open 2026-03 262K n/a n/a $0.15 $0.6 Apache-2.0 workhorse, very cheap
Gemini 3.5 Flash
Google
closed 2026-02 1M 56.0% 55.1% $1.5 $9 Cheap, huge context, fast
Grok Code Fast 1
xAI
closed 2025-08 256K 58.0% n/a $0.2 $1.5 Cheap coding-tuned Grok
Qwen3 235B
Alibaba
open 2025-04 128K 58.0% n/a $0.2 $0.6 Qwen3 flagship reasoner
MiniMax M1
MiniMax
open 2025-06 1M 56.0% n/a $0.2 $1.1 1M-context open MoE
Claude Haiku 4.5
Anthropic
closed 2025-10 200K 55.0% n/a $1 $5 Fast & cheap
GPT-4.1
OpenAI
closed 2025-04 1M 55.0% n/a $2 $8 Long-context workhorse
Llama 4 Maverick
Meta
open 2025-04 1M 55.0% n/a self-host self-host Open, long context
Gemini 2.5 Flash
Google
closed 2025-04 1M 54.0% n/a $0.3 $2.5 Cheap fast workhorse
Grok 3
xAI
closed 2025-02 128K 52.0% n/a $3 $15 xAI's first big reasoning push
Codestral 25.x
Mistral
open 2025-01 256K 51.0% n/a $0.3 $0.9 Lightweight code specialist
Mistral Medium 3
Mistral
closed 2025-05 128K 50.0% n/a $0.4 $2 Mid-tier European model
Claude 3.5 Sonnet
Anthropic
closed 2024-10 200K 49.0% n/a $3 $15 The original agentic coder
DeepSeek R1
DeepSeek
open 2025-01 128K 49.0% n/a $0.55 $2.19 The open-reasoning shock
o3-mini
OpenAI
closed 2025-01 200K 49.0% n/a $1.1 $4.4 Cheap early reasoner
o1
OpenAI
closed 2024-12 200K 48.0% n/a $15 $60 The first reasoning model
Devstral 25.12
Mistral
open 2025-12 262K 72.2% n/a $0.4 $2 Agentic-coding specialist, open weights
Devstral
Mistral
open 2025-05 128K 46.0% n/a self-host self-host Open agentic-coding specialist
Qwen2.5-Coder 32B
Alibaba
open 2024-11 128K 45.0% n/a self-host self-host The open code workhorse
Llama 4 Scout
Meta
open 2025-04 10M 45.0% n/a self-host self-host Open, 10M-token context
Command A
Cohere
closed 2025-03 256K 44.0% n/a $2.5 $10 Enterprise-focused model
DeepSeek V3
DeepSeek
open 2024-12 128K 42.0% n/a $0.27 $1.1 The open MoE breakout
GPT-4.5
OpenAI
closed 2025-02 128K 38.0% n/a $75 $150 Brief, very pricey flagship
Gemini 2.0 Flash
Google
closed 2025-01 1M 35.0% n/a $0.1 $0.4 Cheap, fast, huge context
BugTraceAI-CORE Ultra 27B
BugTraceAI
open 2026-06 4K n/a n/a self-host self-host Offensive-security artifacts (Nuclei, CVE PoCs)

Methodology & caveats. Figures are indicative, compiled from public reports and vendor disclosures as of June 2026, and they move constantly — treat this as a shape-of-the-field snapshot, not a live leaderboard. SWE-bench Verified measures the share of real GitHub issues a model resolves with a passing patch; scores depend heavily on the agent scaffold around the model, so cross-vendor numbers are directional, not head-to-head. Prices are list API rates per million tokens and ignore caching, batching, and volume discounts — per-task cost depends on how many tokens your loop actually burns. Self-host models have no per-token price but a real hardware-and-ops cost (see running models locally). Claude Fable 5 is back on the board: it was withdrawn globally on 2026-06-12 under a US national-security order, but the US lifted the ban on 2026-06-27 and Anthropic is restoring access (vetted US orgs first, broader rollout in the coming weeks — the fallout). Its safeguards-lifted sibling Mythos 5 is listed too, but it stays a restricted rollout with no separate published score, so its SWE-bench shows n/a. The newly previewed GPT-5.6 family (Sol / Terra / Luna) is listed with OpenAI's announced prices but no SWE-bench number yet — OpenAI published only Terminal-Bench, so SWE-bench Verified shows n/a until an independent figure exists; context is inherited from the GPT-5.5 line pending confirmation. GLM-5.2's figure is an early, indicative estimate — Z.ai shipped it without published benchmarks, so treat it as provisional until independently verified. Earlier models' scores are reconstructed from public reports at their time and use representative pricing. The SWE-bench Pro column is the harder, independently-run multi-file variant; it is published for only a handful of frontier models, so most rows show n/a. Several recent rows and figures (prices, context, SWE-bench Verified) are sourced from vibecoding.cz. Always verify against primary sources before a buying decision: SWE-bench, Aider polyglot, LMArena, and each vendor's pricing page.