Qualcomm-AWS, Positron, Broadcom, TPUs: the money is going to inference silicon
On September 8 Qualcomm and Amazon announced a multiyear collaboration on custom silicon for AWS, built for inference rather than training, plus optical interconnect up to 1.6 Tbit/s. Amazon may take up to $60 billion of hardware under the deal, the first western hyperscaler to buy Qualcomm data-centre chips. Two days later Positron, which builds processors specialised for AI inference, closed $875 million at a $5 billion valuation, co-led by NEA, Atreides, Valor, Andra, SemiAnalysis Capital and Jim Clark. The number that ties those together: 8 of the 12 AI-chip funding rounds in 2026 are inference-focused, about 67% of the capital.
Training is a handful of buyers with unlimited budgets. Inference is everyone, every day, on a bill someone actually reads. The money noticed.
The tape
- Qualcomm x AWS (September 8): inference-only custom silicon, 1.6 Tbit/s optics, up to $60 billion in purchases.
- Positron (September 10): $875 million, $5 billion valuation, inference processors.
- Broadcom: AI-chip revenue up 221% to $16.7 billion per quarter, guiding $21.7 billion next. That is the custom-ASIC business, the second engine next to Nvidia GPUs.
- Google TPUs: Oppenheimer projects external TPU sales of about 6.6 GW, roughly $170 billion cumulative by 2028, with around $19 billion of it in 2026.
- Nvidia's response: $3.5 billion into MediaTek plus NVLink Fusion, extending its interconnect into custom accelerators it does not make. Also the $12.93 billion acquisition of Hugging Face with a stated hardware-neutral commitment.
- Agrani Labs (India): a CUDA-compatible inference chip from ex-Intel and AMD engineers, raising about $50 million at a $160 to $200 million valuation.
- China: a planned 100,000-GPU cluster on domestic silicon, and Z.AI reporting that it runs large clusters on domestic Chinese chips with falling inference costs.
Why inference silicon is a different product
A training accelerator is judged on FLOPS and on how many of them you can lash together. An inference accelerator is judged on tokens per second per watt and per dollar at a batch size you actually run. Decode is memory-bound, so the design pressure is bandwidth, on-package memory and interconnect, and the fast optics in the Qualcomm deal are not a footnote; they are the product. The NVFP4 story on Blackwell that I walked through in Blackwell NVFP4 pipeline is Nvidia pursuing the same metric from inside a training-first architecture.
Training bought the GPUs. Inference pays the electricity bill, and the electricity bill is where the new chips are aimed.
Nvidia's moves read as a hedge: NVLink Fusion lets a custom chip sit on Nvidia's fabric, and Hugging Face keeps the distribution channel close. If the accelerator becomes a commodity, own the interconnect and the marketplace.
What it does to your token price
Nothing this quarter. Chips announced in September ship in 2027 and 2028, and the TPU projection runs to 2028. But the direction of the API price list is set by the cost curve of whoever serves the token, and every entry above pushes that curve down for inference specifically. Three things to do with that:
- Model your 2027 unit economics with a declining floor, not a flat one. Hyperscalers with their own inference silicon will price below the GPU-rental floor the moment the chips land, because that is why they built them.
- Keep the routing layer vendor-agnostic. The cheapest token in 2027 may come from a Qualcomm part behind AWS, a TPU behind Google, or a domestic Chinese chip behind an open-weight API. If your stack can only call one, the price drop passes you by. The OpenRouter cost routing setup is the cheapest insurance I know.
- Separate your GPU decisions. If you buy hardware for local inference, the calculus in Hopper vs Blackwell still holds, but the used-GPU market will feel this wave before the price list does.
The honest gap
Almost every figure here is a projection or an announcement. The $60 billion is a ceiling Amazon may reach, not a purchase order; the $170 billion TPU number is an Oppenheimer estimate; Positron's valuation is what investors paid, not what the chips deliver. Only Broadcom's $16.7 billion is a reported quarter. Inference silicon is where the money is going, which is a fact. Whether it is where the performance arrives is a benchmark nobody outside the buyers can run yet.