Kimi K3 on Bedrock: when an open model with caching beats a closed one
On September 18 Moonshot AI's Kimi K3 became generally available on Amazon Bedrock. It's the 2.8 trillion parameter open-weight model with a 1M-token context and native vision. The listing adds something that matters more to my budgets than the parameter count: it is the first open-weight model on Bedrock with explicit prompt caching, and cross-region inference is offered too. Moonshot claims roughly 2.5 times better scaling efficiency than K2. That's their number, I haven't measured it, and the coverage I read doesn't define what "scaling efficiency" covers.
Neither the AWS announcement nor the articles about it that I found give K3 prices on Bedrock, so I'm not going to make any up. What I can do is show how the decision works and leave the price cells for you to fill from the console.
Why caching is most of the story at 1M tokens
Long-context work, a repository in the prompt or an agent carrying 400K tokens of history, resends the same prefix on every call. Without caching you pay the full input price for that prefix each time. With caching you pay a cache-read rate, a fraction of the input price. Past a few dozen calls per session the prefix dominates everything else, and the comparison between models shrinks to two numbers each: fresh input price and cache-read price. Output barely matters for this workload.
The one closed-model reference I do have is Claude Fable 5.1, which I wrote up in the cache read price cut: $10 per million fresh input tokens and $0.25 per million for cache reads.
A worked example, assumptions marked
The workload is assumed: an agent session with a 500,000-token prefix (repository plus instructions), 100 calls, and 2,000 new tokens per call. I count input cost only. I also ignore any cache-write surcharge and any cache expiry, both of which are real and differ between providers.
Fable 5.1 without caching: 100 calls × 502,000 tokens = 50.2M tokens × $10 = $502. With caching: the first call is 0.5M × $10 = $5.00, then 99 calls × 0.5M = 49.5M cache-read tokens × $0.25 = $12.38, plus 100 × 2,000 = 0.2M fresh tokens × $10 = $2.00. Total, about $19.38.
Now K3, with numbers I invented only to show the mechanics. They are not Bedrock prices. Say fresh input is $5 per million, half the Fable figure, and cache reads cost 10% of that, $0.50. The session is 0.5M × $5 = $2.50, plus 49.5M × $0.50 = $24.75, plus 0.2M × $5 = $1.00. That's $28.25, more than the closed model despite a sticker price at half.
| Scenario | Fresh $/M | Cache read $/M | Session input cost |
|---|---|---|---|
| Fable 5.1, no cache | 10 | none | $502.00 |
| Fable 5.1, cached | 10 | 0.25 | $19.38 |
| K3, assumed, 10% cache read | 5 (assumed) | 0.50 (assumed) | $28.25 |
| K3 break-even | 5 (assumed) | 0.32 | $19.38 |
The break-even row comes from a one-line equation. With fresh price $5, the K3 session costs 3.50 + 49.5 × p, where p is the cache-read price per million. Set that equal to $19.38 and p is about $0.32, roughly 6.4% of the fresh price. So under these made-up inputs K3 has to read from cache at under about 6% of its input rate just to tie. The lesson survives whatever the real prices are: compare the cache-read ratio, not the list price.
A model that is half the price per token can still lose, if its cache reads aren't cheap enough.
One more caveat that the arithmetic hides. Every call that lands after the cache has expired is billed as fresh input, and I don't know the cache lifetime Bedrock gives K3. If your agent works in bursts with long gaps, your hit rate is lower than this table assumes, and the winner can flip depending on each provider's expiry window. Measure your real gap distribution before you trust any of this.
Where data governance tips it
The stronger argument for K3 on Bedrock is not price. An AWS customer can run a frontier open-weight model inside its existing account, IAM and logging, without adding a third-party endpoint to the vendor list. For regulated teams that can remove a whole procurement cycle, and it's the only realistic way most of them will use these weights, given how hard the model is to self-host.
I'd check two things before putting that on a slide. Cross-region inference can send requests outside the region you picked, so read the residency terms. And inside your perimeter is not the same as clean provenance: the distillation allegation is still open, and the checklist in the due-diligence piece applies here unchanged.
Also test the window itself. A 1M context that you can fill isn't a 1M context that retrieves accurately, and I haven't run K3 at that length. My suggestion is to take last month's real agent logs, compute the cache-read ratio K3 would have to hit on your traffic, and only then look at the price page.