← all posts
// economics · deepseek

DeepSeek raised prices 2.3x to 4.5x and kept growing: recompute your hosted vs local break-even

DeepSeek raised its API prices by 2.3x to 4.5x, depending on the model, and according to reports still landed at an annualized run-rate of about $1 billion. That is roughly ten times its whole 2025. A financing round of around $7.5 billion is reportedly in the works too. I can't verify any of this from a primary source, and the reports I saw don't give the new per-token price list, so every dollar figure below is my own assumption and labelled as one.

What I do take from it: customers paid more and kept coming. The "cheap Chinese fallback" in your router was never priced at cost, and now the price is moving toward cost. If you wired DeepSeek in as the budget tier after the V4 endpoint retirement, someone in finance is about to ask why that line item grew.

What a 2.3x to 4.5x hike does to a bill

Say your pipeline pushes 400M tokens a month through DeepSeek at a blended price (input plus output, cache-adjusted) of $0.40 per million. That baseline is made up, plug in your own from the invoice. The hike moves the blend to $0.92 at the low end and $1.80 at the high end. Monthly spend goes from $160 to somewhere between $368 and $720.

Small money, honestly. Nobody rebuilds an architecture over $560 a month. But the same multiplier on a team running 20B tokens a month turns $8,000 into $18,400 to $36,000, and that gets a meeting.

Hosted vs local, with the assumptions on the table

The tempting reaction is "fine, I'll run it on a Mac". Let me do the arithmetic before anyone buys hardware. Most inputs below are assumptions. Two come from elsewhere: the M5 Pro price and TDP are the values in the /local-cost table ($3,400, 42 W), and the decode speeds are the ones reported for a heavily quantized Qwen3.8-27B (about 2.4 bits per weight, 8.45 GB) on an M5 Pro: roughly 52 tok/s on code, 17 tok/s on prose. I haven't run that model myself.

assumptions (mine): 36-month amortization, $0.30/kWh, machine on 24/7
fixed/month = 3400/36 + 42W*720h*0.30 = 94.4 + 9.1 = ~$103.5

capacity/month at 100% decode utilization
  25 tok/s -> 64.8M tokens     52 tok/s -> 134.8M tokens

break-even utilization = 103.5 / (hosted $/M * capacity in M)
  hosted $0.92:  25 tok/s -> 174%   52 tok/s -> 83%
  hosted $1.80:  25 tok/s -> 89%    52 tok/s -> 43%

Read it slowly. At a mixed 25 tok/s the Mac never breaks even against the low-end hiked price, because you would need 174% utilization. Against $1.80 it has to be busy 89% of the month, nights included. Only the code-heavy 52 tok/s case looks comfortable, and even there you need 43% to 83% of every hour. At a more realistic 30% utilization, local costs about $5.30 per million tokens at 25 tok/s. That is three to six times the hiked hosted price.

A local box beats a hosted API on price only when it is busy almost all the time, and almost nothing is.

Where this calculation lies to you

Three ways, at least. First, I compared a 27B model at 2.4 bits against a DeepSeek frontier model. Those aren't the same product, so the honest comparison is quality-adjusted, and I haven't run that eval. Second, I counted only decode tokens. Input tokens are prefill, which is far cheaper per token on both sides, and hosted providers discount cached input heavily. Third, batching: a single GPU serving many requests with continuous batching gets an order of magnitude more aggregate throughput than one interactive stream, which changes the capacity line completely. A GPU server, rented or owned, is a different calculation than a laptop, and it is the one to run if your volume is real.

Local still wins on other axes: data that can't leave the building, latency with no network hop, and a cost that does not change when a vendor reprices. Those are worth money, just not the money this spreadsheet measures.

What I would do this week

Pull last month's token counts per model, split into input, cached input and output. Put the hike multipliers on top and see which workloads cross a threshold that matters to you. Then route: cheap and boring steps to the smallest model that passes your eval, cache aggressively, and keep one local model as a fallback for the day a price or an endpoint changes under you. The fallback does not need to save money. It needs to exist and be tested, before the next invoice arrives rather than after. The routing table is the deliverable here, not the spreadsheet.

The question I can't answer from here is whether $1 billion of run-rate after a hike means DeepSeek still has pricing headroom. If it does, this is not the last adjustment.

#deepseek#pricing#local#cost