← all posts
// analysis · inference

Xiaomi HySparse2 cuts 1M-token prefill FLOPs 5.02x. Which part of long-context cost does that actually fix?

Xiaomi's HySparse2 paper (arXiv 2609.26368, posted September 24) claims that, against the Hybrid SWA attention used in MiMo-V2.6, a model built on it needs 5.02x fewer prefill FLOPs at 1M tokens, a 4.5x smaller KV cache, and retrieves better. All three at once. Architecture papers that improve every axis simultaneously make me reach for the fine print, so I spent an evening working out which of those three numbers you would actually feel.

The mechanism, as far as the reports describe it, is two-level KV sharing: layers share key/value state in a hierarchy instead of each layer owning its own cache. Xiaomi says it is aimed at the upcoming MiMo-V3. I haven't read the paper itself, and I haven't run anything, so treat the rest as reasoning about where the money goes, not as a review of the paper.

Two bills that grow differently

Long context is charged twice, and the two charges are different animals. Prefill is compute: you push the whole prompt through the network once, and attention cost grows with the square of length. KV cache is memory: every token you keep leaves key/value vectors in every layer, and that grows linearly but never goes away while the session lives.

Some illustrative numbers. These are my assumptions for a generic dense model with 40 layers, hidden size 4096, 8 KV heads of dimension 128, fp16 cache, 27B parameters. They are not MiMo's real shape.

n = 1,000,000 tokens

attention score+value FLOPs ~ 4 * n^2 * d * layers
  = 4 * 1e12 * 4096 * 40  = ~655 PFLOP
weight matmul FLOPs ~ 2 * params * n
  = 2 * 27e9 * 1e6        = ~54 PFLOP

KV cache = 2 * layers * kv_heads * head_dim * 2 bytes * n
  = 2 * 40 * 8 * 128 * 2 * 1e6 = ~164 GB

Full attention at a million tokens is over ten times the cost of the weights themselves. That is why prefill at 1M feels like waiting for a build. And 164 GB of cache is more than one accelerator holds, before a single weight is loaded. Divide that by 4.5 and you get roughly 36 GB, which suddenly fits next to a quantized model on a big Apple Silicon machine or one datacenter GPU. This is the number I would care about first. Sparse attention that cuts FLOPs but not memory gives you a faster failure to load.

FLOPs are not seconds

Here is where I'd argue with anyone quoting 5x as "5x faster". A FLOPs count says how much arithmetic is required, not how fast hardware performs it. Dense attention runs at high utilization because the access pattern is regular. Sparse and hybrid patterns gather scattered blocks, and utilization drops. If a sparse kernel reaches half the efficiency of the dense one, a 5.02x FLOPs cut turns into something near 2.5x in wall-clock time. I made that ratio up to show the shape of the problem; the real one depends on kernels that, as far as I know, do not exist in public form yet.

Prefill is also only half of a session. Once the prompt is in, decode is bandwidth-bound, and every generated token has to read the cache again. A smaller cache helps there directly, which is another reason I rank the 4.5x above the 5.02x. It is the same lesson as in the KV cache math: the cache sets how many sessions fit and how fast decode runs, and the FLOPs set how long the first token takes.

A FLOPs saving is a claim about arithmetic. A cache saving is a claim about whether the job runs at all.

What that means on a Mac

A second item from the same day is relevant here: Mirai's uzu engine added speculative decoding on Apple M5, reported at nearly 2x over MLX with speculative decoding (MTPLX) and faster than llama.cpp. Good, but it lives on the other side of the split. Speculative decoding speeds up decode, where you are bandwidth-bound and can afford to verify several draft tokens per weight read. It does nothing for a 1M-token prefill, and the gain depends on acceptance rate, which drops on quantized models and odd content. I'd want tokens per second and a quality regression check, not one throughput number. There is more on the prefill/decode split for M5 in this earlier piece.

So if you run agents against huge repositories, the honest reading is this: HySparse2, if it holds up outside the paper, attacks the part of the bill that grows quadratically and the part that decides whether a 1M session fits in memory. The caveat Xiaomi's own framing hints at is a more complicated attention topology and the need to check behavior at the edges of retrieval tasks. Needle-style benchmarks are gentle. I'd test it on the ugly case: a fact buried in the middle of a 900K-token dump of your own logs.

I'd start there when weights land.

#inference#kv-cache#long-context#apple-silicon