Opus 5.5 is cheaper and will break your API calls anyway
Opus 5.5 came out on September 22 at $4 per million input tokens and $20 per million output, down from $5 and $25. Cache reads dropped 60% to $0.20. You read that, you feel good about next month's invoice, and then you reach the line in the release notes saying you can no longer turn thinking off. Cheaper and breaking at the same time. Fun.
The rest of the spec sheet: 1M context, 128k output, always-on adaptive thinking, and a "fast mode" that is a research preview. Anthropic's own comparison against Opus 5 is roughly 40% fewer output tokens and roughly 30% faster generation. Benchmarks from the launch coverage: Terminal-Bench 4.0 at 66.4% (GPT-6 Astra sits at 57.9% in the same comparison), CursorBench 4.0 at 57.8%, Humanity's Last Exam with tools at 67.7%. All vendor-reported, none of it mine. I haven't migrated a production workload yet, so what follows is a plan built from the release facts, not a war story.
Claude Code moved first
Claude Code 2.1.280 makes Opus 5.5 the default model. That's the change most likely to bite you without any code change on your side, because a CI job that runs headless and never names a model just got a different model the day the runner updated. Grep your pipelines for invocations without a model flag and pin them explicitly. Then you upgrade on your schedule, not on the release calendar's. (The headless setup is covered in Claude Code headless in CI if you need the flags.)
Also worth a look: the 1M window at a lower price makes it tempting to stop trimming context. Don't relax yet. The cache read price is cheap, but a million tokens read on every step of a loop is still a million tokens.
Thinking you cannot switch off
This is the change I'd worry about most. On Opus 5.5 thinking is always on and adaptive, meaning the model decides how much to do. Anywhere your code disables thinking to keep a call short and cheap (classification, routing, small extraction jobs, anything with a latency budget) you now get a model that may think anyway.
Two things to do. First, find those code paths. A blunt search is enough to start:
grep -rnE "thinking" src/ --include=*.py --include=*.ts
Second, don't guess at the effect. Replay about 200 real requests from each path and record output tokens and p50/p95 latency before and after. Thinking tokens are billed as output, so the token count is the number that matters. And I haven't checked how the API reacts if you keep sending the old disable parameter: it might be ignored, it might be a hard error. Read the actual response instead of assuming, then delete the dead flags either way. For paths that truly need fast and dumb, a smaller model tier is probably the honest answer, not fighting an adaptive one.
The computer use toolset
Computer use now needs the new tool version, computer_toolset_20260801. The launch notes I saw don't list the schema differences, so open the docs and diff them yourself. Practical advice: put the version string in one constant, run your recorded flows against it in a staging key, and look at click accuracy, not just whether the call succeeds. A toolset with a date in its name will get another date. Make the next upgrade a one-line change.
What the bill actually does
Take an agent-heavy month: 2,000M cache-read tokens, 300M fresh input tokens, 60M output tokens. The Opus 5 cache read price isn't stated anywhere I found, but a 60% cut to $0.20 implies $0.50, which is also exactly a tenth of the old $5 input price, so I'm using it.
| Scenario | Cache reads | Input | Output | Total |
|---|---|---|---|---|
| Opus 5 | $1,000 | $1,500 | $1,500 | $4,000 |
| Opus 5.5, same tokens | $400 | $1,200 | $1,200 | $2,800 |
| Opus 5.5, 40% fewer output | $400 | $1,200 | $720 | $2,320 |
| Opus 5.5, output up 50% | $400 | $1,200 | $1,800 | $3,400 |
A lower price per token only helps if the model doesn't decide to spend more tokens on your behalf.
Even the ugly row is 15% cheaper than before. To lose the whole saving in this mix, output would have to double (the fixed $1,600 of reads and input plus $20 times 60M times 2 lands on $4,000). But that mix is cache-heavy. If your workload is mostly generation with little caching, the only thing that changed is output going from $25 to $20, and then thinking that pushes output up by more than 25% already erases the discount. Know which kind of workload you have before you celebrate.
The 40% fewer output tokens is Anthropic's number, measured against Opus 5, and I'd treat it as a best case until my own logs say otherwise. One more mistake I'd avoid: comparing list prices only. Compare cost per finished task, because a model that thinks more per step but needs fewer steps can win or lose that comparison either way. My plan is a week of shadow traffic on Opus 5.5 with the old model still pinned as a rollback, and a spreadsheet with one column: dollars per merged pull request.