← all posts
// models · models

SWE-2 and the borrowed base: Cognition post-trains Kimi K3 to within a point of frontier

Cognition's new coding model, SWE-2, scores 50.0% on FrontierCode 1.1 Main. That is within a point of Claude Fable 5.1, and Cognition says it costs about 64% less. What caught my eye is not the score but the base: SWE-2 is post-trained from Kimi K3, Moonshot's open-weight model with 2.8 trillion parameters. Cognition did not pre-train anything. They took someone else's weights and pushed them to the frontier with reinforcement learning.

Everything below rests on Cognition's announcement and secondary write-ups. I haven't run SWE-2 myself, and vendor benchmarks are vendor benchmarks.

What Cognition says they built

According to the announcement, SWE-2 beats their own SWE-1.7 and Grok 4.6 on both score and price, and sits a few points behind GPT-6 Astra at roughly a quarter of the cost. It ships inside Devin (Desktop, CLI and Web) and in the Fusion multi-model workflow. The announcement is from about September 10 to 11.

The training claim is the technically interesting part. Cognition describes a single RL run that optimizes several reasoning-effort levels at once (medium, high and max), and it penalizes the cost of the rollout directly. So the reward is not just "did the tests pass". It is closer to "did the tests pass, and how much did you burn getting there".

Cost-aware RL is not a new idea. Doing it at multi-trillion-parameter scale, on a model you did not build, is the new part. The material I worked from calls it the first RL scaling into that regime, which I can't verify independently.

Why a borrowed base is enough

For two years the assumption was that frontier coding needs frontier pre-training, which means a compute bill that only a handful of labs can pay. SWE-2 is a data point against that. If a strong open base exists, the remaining gap in coding is largely a post-training problem: task environments, verifiers, reward design. That is a much smaller bill, and a company sitting on a lot of real coding traces (Cognition has plenty from Devin) is well placed to pay it.

I would not oversell it. Post-training moves the cost/performance frontier. It does not raise the ceiling of what the base model can do, and that is exactly how the source describes it. A model that ties Fable 5.1 on one benchmark can still lose on the odd repo where the base never learned the idiom. I would want to see it on my own code before believing the one-point gap.

There is also a question I keep coming back to: what is the moat of an open base? Moonshot did the expensive part and Cognition harvests the margin. The provenance and licensing worries I covered in Kimi K3 distillation due diligence apply here too. And the practical side, that K3 is open but you will not run 2.8T parameters on your desk, is in Kimi K3: open but unrunnable. SWE-2 is the version of that story where someone else hosts it for you.

If the base is open, frontier coding turns into a post-training contest, and that is a contest a smaller player can enter.

What 64% does to your bill

Take a made-up but plausible number. Say a team's agents spend 10,000 dollars a month on a frontier model. At 64% off, the same volume costs 3,600. You save 6,400 a month, or 76,800 a year. Flip it around: for 10,000 you now buy 2.8 times the agent work (1 / 0.36). Whether that is a saving or a bigger fleet is a management choice, and given how agent usage behaves, I would bet on the bigger fleet.

Two caveats on that arithmetic. First, price per token is not price per solved task. A cheaper model that needs more retries or longer runs can eat the discount, which is precisely why the cost-of-rollout penalty in training matters. Second, I am reading the 64% as relative to Fable 5.1, but the material I had does not spell out the baseline, so check Cognition's own table before quoting it.

Multiple effort levels from one run also change how you would route. Instead of maintaining a small model for easy tickets and a big one for hard ones, you could send everything to one model and dial effort per task. Less plumbing, one thing to evaluate. I like that more than the headline number.

The pressure on pricing is the real story. When a smaller vendor can get to within a point at a third of the price, the frontier labs' per-token margins on coding are not safe, and I expect the next round of price moves to show it. If you pay for agents at scale, this is the week to run your own benchmark set against a cheaper contender, using your repos and your failing tests, not theirs.

#models#cost#agents#economics