The LLM coding benchmark
Every coding model that mattered, from the dawn of the agentic-coding era (late 2024) to today: when it shipped, how big a window, how well it closes real GitHub issues, and what it costs. Click any column to sort — try Released for the full timeline.
| Model | Weights | Released | Context | SWE-bench Verified | SWE-bench Pro | Input /M | Output /M | Best for |
|---|---|---|---|---|---|---|---|---|
Claude Opus 5 Anthropic |
closed | 2026-07 | 1M | n/a | n/a | $5 | $25 | Newest Anthropic flagship; 1M context; no published SWE-bench Verified run yet |
Claude Opus 5 (fast mode) Anthropic |
closed | 2026-07 | 1M | n/a | n/a | $10 | $50 | Low-latency Opus 5 tier, 2x standard price; no published Verified run yet |
Claude Mythos 5 Anthropic |
closed | 2026-06 | 1M | n/a | n/a | $10 | $50 | Safeguards-lifted Mythos tier; restricted rollout |
Claude Fable 5 Anthropic |
closed | 2026-06 | 1M | 95.0% | 80.3% | $10 | $50 | Mythos-class frontier (restored after US ban) |
GPT-5.6 Sol OpenAI |
closed | 2026-06 | 400K | 96.2% | 64.6% | $5 | $30 | Flagship; Terminal-Bench leader; Verified score from vals.ai independent run (OpenAI retired the benchmark) |
GPT-5.6 Terra OpenAI |
closed | 2026-06 | 400K | n/a | 63.4% | $2.5 | $15 | Lower-cost, ~GPT-5.5 class; no published Verified run |
GPT-5.6 Luna OpenAI |
closed | 2026-06 | 400K | 93.0% | 62.7% | $1 | $6 | Fastest, cheapest GPT-5.6 tier; Verified score from vals.ai independent run |
Grok 4.5 xAI |
closed | 2026-07 | 500K | n/a | 64.7% | $2 | $6 | Newest xAI flagship; SWE-bench Pro 64.7%, no Verified score |
Gemini 3.6 Flash Google |
closed | 2026-07 | 1M | n/a | 58.7% | $1.5 | $7.5 | Agent-tuned Flash, ~17% fewer output tokens; SWE-bench Pro 58.7%, no Verified score |
Claude Sonnet 5 Anthropic |
closed | 2026-06 | 1M | 72.7% | n/a | $3 | $15 | Strong agentic build follow-through, pricier per task than 4.6 |
Claude Opus 4.8 Anthropic |
closed | 2026-05 | 1M | 78.0% | 69.2% | $5 | $25 | Frontier agentic coding |
Claude Opus 4.8 (fast mode) Anthropic |
closed | 2026-05 | 1M | 78.0% | n/a | $10 | $50 | Preview low-latency tier, 2x standard Opus 4.8 price |
GLM-5.2 Zhipu AI |
open | 2026-06 | 200K | 78.0% | n/a | $1 | $3 | Latest open frontier (indicative) |
GLM-5 Zhipu AI |
open | 2026-04 | 200K | 77.8% | n/a | $1 | $3 | Strong open coder (Z.ai) |
GPT-5.5 Codex OpenAI |
closed | 2026-04 | 400K | 77.0% | n/a | $5 | $30 | Tuned for the Codex agent |
Claude Opus 4.7 Anthropic |
closed | 2026-03 | 1M | 77.0% | n/a | $5 | $25 | Previous Opus flagship |
Claude Opus 4.6 Anthropic |
closed | 2025-12 | 1M | 77.0% | n/a | $5 | $25 | Adaptive-thinking Opus |
GPT-5.5 OpenAI |
closed | 2026-04 | 400K | 76.0% | 58.6% | $5 | $30 | OpenAI flagship |
Claude Opus 4.5 Anthropic |
closed | 2025-11 | 200K | 76.0% | n/a | $5 | $25 | First effort-dial Opus |
Claude Opus 4.1 Anthropic |
closed | 2025-08 | 200K | 74.0% | n/a | $15 | $75 | Older frontier Opus |
GPT-5.4 OpenAI |
closed | 2026-03 | 400K | 74.0% | n/a | $2.5 | $15 | Cheaper GPT-5 tier |
Qwen3.5 Plus Alibaba |
closed | 2026-02 | 1M | 76.4% | n/a | $0.4 | $2.4 | Cheap 1M-context hosted coder |
Grok 4.20 xAI |
closed | 2026-03 | 2M | 76.7% | n/a | $2 | $6 | 2M-context Grok, superseded by 4.3/4.5 |
GPT-5.4 Mini OpenAI |
closed | 2026-03 | 400K | 48.2% | n/a | $0.75 | $4.5 | Cheap GPT-5.4 tier, 128k output |
Claude Sonnet 4.6 Anthropic |
closed | 2026-01 | 1M | 73.0% | n/a | $3 | $15 | Best price/perf for agents |
GPT-5 Codex OpenAI |
closed | 2025-09 | 400K | 73.0% | n/a | $1.25 | $10 | First Codex-tuned GPT-5 |
Qwen3.7-Max Alibaba |
closed | 2026-03 | 1M | 72.5% | n/a | $2.5 | $7.5 | 1M-context agent model |
Claude Sonnet 4.5 Anthropic |
closed | 2025-09 | 200K | 72.0% | n/a | $3 | $15 | Strong cheaper agent model |
Claude Opus 4 Anthropic |
closed | 2025-05 | 200K | 72.0% | n/a | $15 | $75 | First Claude 4 Opus |
Claude Sonnet 4 Anthropic |
closed | 2025-05 | 200K | 71.0% | n/a | $3 | $15 | First Claude 4 Sonnet |
GPT-5 OpenAI |
closed | 2025-08 | 400K | 72.0% | n/a | $1.25 | $10 | The GPT-5 reasoning leap |
Gemini 3.1 Pro Google |
closed | 2026-02 | 1M | 72.0% | n/a | $2 | $12 | 1M context, tiered pricing |
Grok 4.3 xAI |
closed | 2026-05 | 256K | 71.0% | n/a | $1.25 | $2.5 | Cheap frontier, fast |
Grok Build 0.1 xAI |
closed | 2026-05 | 256K | 70.8% | n/a | $1 | $2 | Coding-agent-tuned Grok, preview |
Gemini 3 Pro Google |
closed | 2025-11 | 1M | 71.0% | n/a | $2 | $12 | The Gemini 3 jump |
o3 OpenAI |
closed | 2025-04 | 200K | 69.0% | n/a | $2 | $8 | The o-series reasoning peak |
Kimi K2.6 Moonshot AI |
open | 2026-04 | 256K | 68.0% | n/a | $0.95 | $4 | Open agentic coder, 1T MoE |
DeepSeek V4 Pro DeepSeek |
open | 2026-04 | 1M | 80.6% | n/a | $0.44 | $0.87 | Highest open-weight SWE-bench Verified score, very cheap |
o4-mini OpenAI |
closed | 2025-04 | 200K | 68.0% | n/a | $1.1 | $4.4 | Cheap strong reasoner |
Grok 4 xAI |
closed | 2025-07 | 256K | 68.0% | n/a | $3 | $15 | The frontier Grok |
GLM-4.6 Zhipu AI |
open | 2025-10 | 200K | 68.0% | n/a | $0.6 | $2.2 | Stronger open GLM coder |
Qwen3 Max Alibaba |
closed | 2025-09 | 256K | 67.0% | n/a | $0.78 | $3.9 | First Qwen3 Max |
Qwen3-Coder 480B Alibaba |
open | 2025-07 | 256K | 66.0% | n/a | self-host | self-host | Top open code specialist |
DeepSeek V3.1 DeepSeek |
open | 2025-08 | 128K | 66.0% | n/a | $0.27 | $1.1 | Open frontier (V3 line) |
Kimi K2 Moonshot AI |
open | 2025-07 | 256K | 65.8% | n/a | $0.55 | $2.2 | The original K2 |
Gemini 2.5 Pro Google |
closed | 2025-03 | 1M | 64.0% | n/a | $1.25 | $10 | The long-context breakout |
GLM-4.5 Zhipu AI |
open | 2025-07 | 128K | 64.0% | n/a | $0.6 | $2.2 | The GLM open-coder breakout |
MiniMax M3 MiniMax |
open | 2026-06 | 200K | 63.0% | n/a | $0.3 | $1.2 | Cheap, agentic, open |
Claude Sonnet 3.7 Anthropic |
closed | 2025-02 | 200K | 62.0% | n/a | $3 | $15 | The agentic-coding breakout |
DeepSeek V4 Flash DeepSeek |
open | 2026-04 | 1M | 60.0% | n/a | $0.14 | $0.28 | Cheapest frontier-class |
MiniMax M2 MiniMax |
open | 2025-10 | 200K | 60.0% | n/a | $0.3 | $1.2 | Earlier M-series open MoE |
Mistral Medium 3.5 Mistral |
closed | 2026-04 | 262K | 77.6% | n/a | $1.5 | $7.5 | Mistral's strongest current model |
Mistral Large 3 Mistral |
closed | 2026-04 | 256K | 58.0% | n/a | $0.5 | $1.5 | Cheap European flagship |
Mistral Small 26.03 Mistral |
open | 2026-03 | 262K | n/a | n/a | $0.15 | $0.6 | Apache-2.0 workhorse, very cheap |
Gemini 3.5 Flash Google |
closed | 2026-02 | 1M | 56.0% | 55.1% | $1.5 | $9 | Cheap, huge context, fast |
Grok Code Fast 1 xAI |
closed | 2025-08 | 256K | 58.0% | n/a | $0.2 | $1.5 | Cheap coding-tuned Grok |
Qwen3 235B Alibaba |
open | 2025-04 | 128K | 58.0% | n/a | $0.2 | $0.6 | Qwen3 flagship reasoner |
MiniMax M1 MiniMax |
open | 2025-06 | 1M | 56.0% | n/a | $0.2 | $1.1 | 1M-context open MoE |
Claude Haiku 4.5 Anthropic |
closed | 2025-10 | 200K | 55.0% | n/a | $1 | $5 | Fast & cheap |
GPT-4.1 OpenAI |
closed | 2025-04 | 1M | 55.0% | n/a | $2 | $8 | Long-context workhorse |
Llama 4 Maverick Meta |
open | 2025-04 | 1M | 55.0% | n/a | self-host | self-host | Open, long context |
Gemini 2.5 Flash Google |
closed | 2025-04 | 1M | 54.0% | n/a | $0.3 | $2.5 | Cheap fast workhorse |
Grok 3 xAI |
closed | 2025-02 | 128K | 52.0% | n/a | $3 | $15 | xAI's first big reasoning push |
Codestral 25.x Mistral |
open | 2025-01 | 256K | 51.0% | n/a | $0.3 | $0.9 | Lightweight code specialist |
Mistral Medium 3 Mistral |
closed | 2025-05 | 128K | 50.0% | n/a | $0.4 | $2 | Mid-tier European model |
Claude 3.5 Sonnet Anthropic |
closed | 2024-10 | 200K | 49.0% | n/a | $3 | $15 | The original agentic coder |
DeepSeek R1 DeepSeek |
open | 2025-01 | 128K | 49.0% | n/a | $0.55 | $2.19 | The open-reasoning shock |
o3-mini OpenAI |
closed | 2025-01 | 200K | 49.0% | n/a | $1.1 | $4.4 | Cheap early reasoner |
o1 OpenAI |
closed | 2024-12 | 200K | 48.0% | n/a | $15 | $60 | The first reasoning model |
Devstral 25.12 Mistral |
open | 2025-12 | 262K | 72.2% | n/a | $0.4 | $2 | Agentic-coding specialist, open weights |
Devstral Mistral |
open | 2025-05 | 128K | 46.0% | n/a | self-host | self-host | Open agentic-coding specialist |
Qwen2.5-Coder 32B Alibaba |
open | 2024-11 | 128K | 45.0% | n/a | self-host | self-host | The open code workhorse |
Llama 4 Scout Meta |
open | 2025-04 | 10M | 45.0% | n/a | self-host | self-host | Open, 10M-token context |
Command A Cohere |
closed | 2025-03 | 256K | 44.0% | n/a | $2.5 | $10 | Enterprise-focused model |
DeepSeek V3 DeepSeek |
open | 2024-12 | 128K | 42.0% | n/a | $0.27 | $1.1 | The open MoE breakout |
GPT-4.5 OpenAI |
closed | 2025-02 | 128K | 38.0% | n/a | $75 | $150 | Brief, very pricey flagship |
Gemini 2.0 Flash Google |
closed | 2025-01 | 1M | 35.0% | n/a | $0.1 | $0.4 | Cheap, fast, huge context |
BugTraceAI-CORE Ultra 27B BugTraceAI |
open | 2026-06 | 4K | n/a | n/a | self-host | self-host | Offensive-security artifacts (Nuclei, CVE PoCs) |
Methodology & caveats. Figures are indicative, compiled from public reports and vendor disclosures as of June 2026, and they move constantly — treat this as a shape-of-the-field snapshot, not a live leaderboard. SWE-bench Verified measures the share of real GitHub issues a model resolves with a passing patch; scores depend heavily on the agent scaffold around the model, so cross-vendor numbers are directional, not head-to-head. Prices are list API rates per million tokens and ignore caching, batching, and volume discounts — per-task cost depends on how many tokens your loop actually burns. Self-host models have no per-token price but a real hardware-and-ops cost (see running models locally). Claude Fable 5 is back on the board: it was withdrawn globally on 2026-06-12 under a US national-security order, but the US lifted the ban on 2026-06-27 and Anthropic is restoring access (vetted US orgs first, broader rollout in the coming weeks — the fallout). Its safeguards-lifted sibling Mythos 5 is listed too, but it stays a restricted rollout with no separate published score, so its SWE-bench shows n/a. The newly previewed GPT-5.6 family (Sol / Terra / Luna) is listed with OpenAI's announced prices but no SWE-bench number yet — OpenAI published only Terminal-Bench, so SWE-bench Verified shows n/a until an independent figure exists; context is inherited from the GPT-5.5 line pending confirmation. GLM-5.2's figure is an early, indicative estimate — Z.ai shipped it without published benchmarks, so treat it as provisional until independently verified. Earlier models' scores are reconstructed from public reports at their time and use representative pricing. The SWE-bench Pro column is the harder, independently-run multi-file variant; it is published for only a handful of frontier models, so most rows show n/a. Several recent rows and figures (prices, context, SWE-bench Verified) are sourced from vibecoding.cz. Always verify against primary sources before a buying decision: SWE-bench, Aider polyglot, LMArena, and each vendor's pricing page.