Sonnet 5.5: same price, 30% faster, and a setting your thinking-off workloads must adopt
Anthropic released Claude Sonnet 5.5 on September 28. The price stayed at $2 per million input tokens and $10 per million output, with cache reads at $0.20 and cache writes at $2.50. Generation is 30% or more faster than Sonnet 5, and it's available on the Claude Platform, AWS, Google Cloud and Azure. The headline numbers are Terminal-Bench 4.0 at 70.6% (Sonnet 5 scored 10.3% on it), OSWorld 2.1 at 80.1%, and FrontierCode 1.1 at 46.2% in Max and 52.1% in Xhigh.
Everyone will quote the benchmarks. I'd start with the speed, because it's the one number here that transfers to your workload without an asterisk.
Same price per token, a different price per run
Token price didn't move, so a token-for-token workload costs the same. Where speed pays is everything billed or budgeted in wall-clock time. An agent run that spends 200 seconds generating takes about 154 seconds if "30% faster" means 1.3 times the throughput, or 140 seconds if it means 30% less time. I don't know which reading Anthropic intends, so the honest range is 140 to 154. Either way, a fifth to a third of the generation wait disappears.
That matters more than it sounds in three places. Sandboxes and VMs that bill by the minute run shorter. A human waiting on a coding agent hits fewer moments where they alt-tab and lose the thread. And latency budgets in production (the ones your product manager set at 30 seconds end to end) suddenly have room for one more tool call or one retry, which is often the difference between a run that recovers from a failed command and one that gives up.
Cache prices are the quiet half. A read at $0.20 is a tenth of the base input price, and a write at $2.50 is 1.25 times it. In a long agent loop where the same 60K-token context is replayed on each step, the cache read is doing most of the financial work, exactly as in the earlier Sonnet 5 pricing story. Nothing new to do here, just keep caching on and check your hit rate after migrating, because a changed system prompt or tool list during the move silently resets the cache and you pay the $2.50 write again on every run until it settles.
A seven-fold jump deserves suspicion
Terminal-Bench 4.0 going from 10.3% to 70.6% is a big number, and I'd resist turning it into a claim about your tasks. The 4.0 version is a new and tougher test, so the low Sonnet 5 score says something about the benchmark as much as about the predecessor. And a jump that large usually means the model crossed a capability threshold on a specific kind of task (long terminal sessions with recovery from errors), not that it's seven times better at everything.
There is a useful way to read it anyway. If a failed attempt costs about the same as a successful one, cost per solved task scales with one over the pass rate: 1 / 0.706 is about 1.4 attempts per success, 1 / 0.103 is about 9.7. That's the shape of the argument for agent economics. It's also a fantasy in one respect, because failed runs are correlated (the same bad task fails again), so treat the ratio as a ceiling on the improvement and measure your own.
Speed is a property of the model. A benchmark score is a property of somebody else's tasks.
What the migration asks of you
The one required change I know of: workloads that run Sonnet with thinking turned off must move to the new between_tools setting before switching to 5.5, or they can fail at the transition. I haven't seen the exact parameter shape yet, so I won't invent it here. Read the release notes and the API reference for the accepted values before touching code.
What I can tell you is how I'd find the affected calls. Search for every place that constructs a request to a Sonnet model and look at how thinking is configured:
grep -rn "thinking" --include="*.py" --include="*.ts" --include="*.kt" .
grep -rn "sonnet" config/ infra/ 2>/dev/null
Anything that pins a Sonnet model ID and doesn't enable thinking is a candidate. Plugins, CI bots, client integrations that somebody wrote in March and nobody has opened since. If you're also moving Opus, the Opus 5.5 migration checklist covers the thinking, computer-use and toolset differences on that side.
After the change, rerun a fixed task set on both models and record three numbers per task: success, total tokens, wall-clock time. Cost per task is tokens times price, and the wall-clock column is where the 30% should show up. If it doesn't, the cause is probably tool latency, not generation, and the model can't fix that for you. Slow shell commands, a cold test suite and a rate-limited API are all outside the 30%.