61% of OpenRouter tokens are Chinese models, and your invoice will never show it
The number going around is that Chinese models account for about 61% of all tokens routed through OpenRouter. CNBC ran it on September 26, and the breakdown is what makes it worth reading: Xiaomi MiMo at roughly 21% of volume, DeepSeek at 17.6%, Anthropic at 15.4%, Google down from about 37% a year ago to about 13%, and Meta's Llama below 1%. Committees in the US House have started an inquiry. What exactly they are asking for I haven't seen, and the coverage I read doesn't say.
I keep seeing this quoted as market share. It isn't. It's token share on one aggregator, and tokens are a bad unit for anything involving money.
Sixty-one percent of tokens, a rounding error of revenue
The best evidence is in the same reporting. DeepSeek on Vercel went from under 1% of tokens to about 17%, while bringing in roughly 1% of revenue. Volume up, margin nowhere. Cheap open weights take the tonnage and leave the money, and they drag everyone's inference price down while doing it.
A toy calculation with made-up round numbers (not real price lists): say a cheap open-weight model costs $0.40 per million output tokens and a frontier model $10. Split the traffic 61/39. The cheap side contributes 0.61 × 0.40 = $0.244 per million blended tokens, the frontier side 0.39 × 10 = $3.90. Spend share of the cheap models: 0.244 / 4.144, about 5.9%. Sixty-one percent of the tokens, under six percent of the bill, and I didn't even use an extreme price gap.
There is a second bias in the data. OpenRouter is where people go when they want to shop on price, by construction. Nobody needs a router if they are happy with one vendor contract, so a bank on a hyperscaler enterprise agreement doesn't appear in this chart at all. Coding agents, hobby projects and free tiers are heavily represented (MiMo's share of coding traffic is a hair higher than its overall share, about 22%). The 61% describes the price-sensitive edge of the market, and that's a legitimate population, just not the market.
What it does to a routing table
If you already read the OpenRouter cost routing piece, none of this changes the mechanics. It changes the default. A year ago the cheap tier in a routing table was a small model from a Western lab. Now the cheap tier is very often a large open-weight model that is close enough on bulk work: extraction, classification, summarising tickets, first-pass code review.
So route bulk there, per task class, with an eval you wrote yourself. Not per leaderboard position. The cost that matters is the cost of a wrong answer: a misclassified support ticket costs cents, a wrong migration script costs an engineer's afternoon, and that asymmetry decides which steps deserve the frontier price. I'd keep the expensive model on the steps where a failure is discovered late.
Weights, endpoints and where the bytes go
"Chinese model" bundles three different things, and compliance cares about only some of them. One is weights trained by a Chinese lab. Another is an endpoint operated by that lab, which means your prompts travel to that jurisdiction. The third is the same weights served by a US or EU host, or by you, on hardware you own. The legal exposure of the second is nothing like the third, and a model name in a config file doesn't tell you which one you have.
On OpenRouter you can pin this in the request instead of hoping. Roughly (check the current docs, I haven't re-verified the field names this week):
"provider": {
"only": ["your-approved-host"],
"allow_fallbacks": false,
"data_collection": "deny"
}
The allow_fallbacks line is the one people forget. Without it, a request that fails on your approved host can quietly land on a provider you never reviewed, and no error tells you so.
Residency isn't the whole story either. Even self-hosted weights raise a supply-chain question about behavior you can't fully audit: tool-call habits, refusal patterns, quirks from fine-tuning data. I have no evidence of anything malicious in these models, and I'm not implying it. But my guess is that "model origin" shows up as a line in procurement questionnaires for regulated industries well before any statute names it, and the House inquiry makes that sooner rather than later. That part is opinion, and I'd happily be wrong about the timing.
A token count tells you what people are running. It says nothing about what they are allowed to run.
The question I'd bring to a client isn't whether to move part of the stack to Qwen, DeepSeek or MiMo. It's which of your tasks can tolerate a model you can't fully audit, on an endpoint you can pin. Pull the ten prompts that burn the most tokens in your own logs and start there.