← all posts
// analysis · deepseek

DeepSeek in mid-2026: the open frontier that keeps undercutting everyone

I've written up GLM, Kimi, Qwen, and Mistral on this site and somehow never given DeepSeek its own post, which in hindsight is an odd gap given it's the lab that keeps landing punches nobody expects. V4 Pro's SWE-bench Verified score, 80.6%, is quietly the most interesting number on the benchmark right now: it's the highest open-weight score on the board, ahead of Opus 4.8, GLM-5.2, Mistral Medium 3.5, and everything closed except the very top of the frontier (Fable 5 and the GPT-5.6 line). An open-weight, MIT-licensed model just beat Opus-class closed models on the one yardstick everyone actually checks. That's not a footnote.

The lineup

V4 Pro is the flagship: a 1.6T-parameter MoE with 49B active per token, 1M context, MIT license, priced at $0.44 in / $0.87 out per million tokens. V4 Flash is the same architecture family scaled down (284B total, 13B active, same 1M context, $0.14 / $0.28), the tier you reach for when Pro's quality is overkill and the bill still needs to be boring. Behind those two sits the line that got DeepSeek onto everyone's radar in the first place: R1, the January 2025 reasoning model that proved an open lab could ship competitive chain-of-thought weights at a fraction of what OpenAI and Anthropic were charging, and V3 and V3.1, the dense-MoE releases that came before V4 tightened the whole lineup.

The price war it keeps winning

Line V4 Pro's $0.44 / $0.87 up against the models it now matches or beats on pricing: Opus 4.8 runs $5 / $25, GPT-5.6 Sol runs $5 / $30, Fable 5 runs $10 / $50. That's an order of magnitude on input and closer to two on output, against a model that's currently sitting in the top tier of the whole leaderboard. DeepSeek didn't get cheap by getting worse; it got cheap while closing most of the gap to the top.

An open-weight, MIT-licensed model beating Opus-class closed models on the one benchmark everyone actually checks is not a footnote.

Self-hosting: technically open, practically heavy

MIT license means you can run it yourself, and that's where the honest math shows up. The 2×H100 96GB sizing table already says no to DeepSeek's older 671B R1/V3 line: 350+ GB even at an aggressive 4-bit quant, nowhere near a 188 GB pair. V4 Pro's 1.6T total parameters make that a louder no, not a smaller one: this is a 4-to-8-GPU proposition or a CPU-offload science project, not a weekend box, whatever the per-token API price suggests about how "light" the model is. V4 Flash is the one actually within reach of a serious multi-GPU workstation, in the same rough footprint class as the other 200-300B MoEs that do fit with room for real concurrency. If your team's GPU budget lives on /local-cost, Flash is the DeepSeek tier to model against it. Pro is a hosted-API decision for almost everyone, not a self-hosting one.

Where it actually fits

  • API-first, cost-sensitive teams: V4 Pro through OpenRouter or DeepSeek's own API is the highest-signal cheap tier on the board right now, full stop.
  • Team GPU box, own the weights: V4 Flash is the realistic self-host target; V4 Pro isn't, regardless of what the price tag implies.
  • Laptop or single consumer GPU: neither fits directly: squeeze what the local tier can actually do with a smaller model first, and route the rest to V4 Pro's API price rather than fighting a quant that won't be kind to quality.
  • Regulated or data-residency-sensitive work: MIT weights only solve the "can I run it myself" question, not the "should this vendor's model touch this data" one. That's a policy conversation, not a pricing one, and it's the same one export controls already put on the table for open-weight Chinese labs generally.

The honest gap

DeepSeek's numbers come from the vendor, same as everyone else's, and I hold them to the standard I hold every lab to on this site: V4 Pro's SWE-bench Verified figure is published, so it's on the table; V4 Flash doesn't have a separately reported one, so its row stays an estimate until that changes. The open-weights thesis this site keeps making (capability you can run anywhere, price pressure that drags the whole market down) isn't hypothetical anymore. DeepSeek is currently the cleanest proof of it running.

#deepseek#analysis#open-weights#cost