GPT-6 Sol and Luna: changing reasoning effort without paying for the cache again
The price cut is the headline, but it isn't the line in the GPT-6 Sol and Luna announcement that I'd build on. That line is a small one: you can change the reasoning effort between calls without invalidating the prompt cache. For anyone running an agent loop, that is worth more than another 50% off.
The numbers first. OpenAI released two cheaper GPT-6 tiers on September 22. Sol costs $2 per million input tokens and $10 output, half of GPT-5.6. Luna costs $0.10 and $0.50. Cached tokens get up to a 90% discount. Sol scores 68.8% on DeepSWE v1.1, which the coverage puts at the level of Anthropic's Fable for around 20% of the price, and Luna is said to get close to Opus 5 and Fable 5 at medium effort. Those are launch claims. I haven't run either model. (The tier I wrote about earlier, in GPT-6 Astra and price convergence, now looks like the expensive end of a much wider family.)
Why a cache cares about effort at all
A prompt cache stores the processed prefix of your request. If the next call starts with exactly the same tokens, you skip recomputing them and pay the cached rate. In an agent loop the prefix is huge and grows every step: system prompt, tool definitions, the whole history of file reads and command output. That is why the 90% matters, and why anything that breaks the prefix hurts.
Reasoning effort is a setting, not text, so you'd hope it stays out of that prefix. The fact that OpenAI advertises it as a feature suggests it wasn't a given. I don't know how OpenAI implements this and won't pretend to. My guess is that effort only affects how the model generates after the prefix, so there is nothing to recompute. The observable claim is enough for the design argument: turn the dial per step, keep the discount.
What a loop looks like when the dial is free
Most steps in an agent run are boring. Read a file, run the tests, look at the diff. A few are hard: choosing the approach, diagnosing the failure that isn't obvious. Fixed high effort pays for deep thinking on the boring ones. Fixed low effort saves money and then fails on the hard ones. With a free dial you route by step type: low for tool plumbing, high for planning and debugging.
Here is a small calculation. Assume 30 steps, a cached prefix averaging 80k tokens, 2k fresh input tokens per step, on Sol at $2 input and $10 output, cached reads at the best case of $0.20 per million. Say a hard step produces 4,000 output tokens including reasoning, and an easy step 400. Those output sizes are my assumptions, not measurements.
| Strategy | Cached reads | Fresh input | Output | Total per run |
|---|---|---|---|---|
| High effort on all 30 steps | $0.48 | $0.12 | $1.20 | $1.80 |
| High on 5 hard steps, low on 25 | $0.48 | $0.12 | $0.30 | $0.90 |
Half the cost, same prefix reuse. Now break the mechanism on purpose: suppose each effort switch forced a cold cache and you had to pay full input price for the 80k prefix again. That is 80,000 tokens at $2 per million, minus what you would have paid at the cached rate, so about $0.14 per switch. The saving from routing is $0.90 per run, so you can afford about six switches (0.90 divided by 0.144). A run where hard steps are scattered, say five separate hard steps needing ten switches up and down, would cost around $2.34 and end up worse than never routing at all.
The interesting property of cache-stable effort is that it removes a reason not to be smart about routing.
Where I'd still be careful
Cached price is "up to" 90%, and cache entries expire, so the idle time between steps matters as much as the setting. A loop that waits on a slow build for longer than the cache lifetime pays the cold price regardless. The other unknown is quality: does a model dropping to low effort mid-conversation behave well when its own earlier high-effort reasoning is sitting in the history? I'd test that before trusting the routing table, mostly to see whether it turns oddly terse after its own long chains.
Then there is which tier. Luna at $0.10 and $0.50 is cheap enough that the honest first experiment is not effort routing at all but running your easy steps on Luna and your hard ones on Sol. That does invalidate the cache, since it is a different model, and my arithmetic above says a switch costs about $0.14 on Sol-sized prefixes. On Luna the prefix is far cheaper, so I suspect model routing wins for coarse phases and effort routing wins for fine ones. Measure both on your own traces, with your own step mix, before believing my numbers or OpenAI's.