← all posts
// hardware · vllm

vLLM and DeepSeek V4.1: 288 bytes per token on AMD gfx950

I opened the vLLM pull request expecting a headline about AMD catching up, and found a much more useful number buried in the description: 288 bytes per token for the compressed KV cache, against 584 bytes uncompressed. That is 296 bytes saved per token, a cut of 49.3 percent, or roughly 2.03x. Half the cache, for a model that already spends its architecture budget on making the cache small.

What the pull request actually does

PR #57463 in vllm-project/vllm adds the NVFP4 compressed KV cache for DeepSeek V4.1 on AMD gfx950, the MI355X generation. Until now that path was NVIDIA only, for two reasons named in the description: the quantization step was written in PTX assembly, and the read side was gated on SM100. The PR swaps the PTX in the fp32x2 to fp4x2 conversion for the AMD instruction v_cvt_scalef32_pk_fp4_f32, which handles the low nibble first and applies an E8M0 scale. On the read side it adds a gfx950 sparse decode tile built on v_cvt_scalef32_pk_f32_fp4, with the e4m3 scales applied separately.

One caveat before anything else: when I fetched it, the PR was open and flagged as needing a rebase. The morning digest I started from described the wider DeepSeek V4.1 kernel work (MegaAttention, Sparse Logits Indexer and friends) as integrated on NVIDIA, but I could only confirm the ROCm piece, and only as a proposal. Nothing here is a release note. If you run vLLM in production, this is something to watch, not something to pin.

A halved KV cache that buys 2.5 percent more concurrency tells you the cache was never the bottleneck you assumed.

The arithmetic of 288 versus 584

The PR does not say whether those byte counts are per layer or per token across the whole model, and I did not dig through the code to find out, so treat the absolute numbers as a ratio only. The ratio is clean enough. If the figures are per token for the full stack, a 128k context costs 128,000 x 584 = about 74.8 MB uncompressed and 128,000 x 288 = about 36.9 MB compressed. That is an illustration of the scaling, not a measured footprint. Even if I am off by a large constant factor, the halving carries over.

Now the part that does not add up at first glance. The PR reports a concurrency gain of only 2.5 percent from the compression, with a possible 24 percent if a sliding-window optimization lands. If the cache shrank by half, why did capacity barely move? My reading is that in the tested configuration the KV cache was not the binding term. Weights, activations and the sliding-window layers that keep a fixed-size window regardless of context all eat memory that compression does not touch. I cannot verify that from the description, it is an inference, and the 24 percent figure is explicitly a projection.

Accuracy and latency, read carefully

The reported cost is small. Decode latency stays within 1 percent across 1 to 64 decode tokens, and GSM8K lands at 90.30 percent on the 584-byte record against 89.92 percent on the compressed one. That is a 0.38 point drop. On a single benchmark run, with no spread reported, I would not call that either free or costly. GSM8K is also a short-context task, so it says little about what 4-bit keys and values do to recall at 100k tokens, which is exactly where you would want the compression.

The other detail I like: the PR says it was developed with Claude Code assistance. Kernel work that maps one vendor's conversion instructions onto another's is the kind of tedious, well-specified translation where that makes sense, and the latency and accuracy tables are what a reviewer should lean on instead of the authorship note.

Why I care if I never buy an MI355X

Two reasons. First, it shows NVFP4 as a format, not only an NVIDIA feature: the scale layout (E8M0 on the write side, e4m3 on the read side) has to be reproduced faithfully by another vendor's hardware, and that is a good test of whether the format is actually portable. I covered the NVIDIA side of the pipeline in Blackwell NVFP4 pipeline, and this is the first time I have seen the KV cache specifically move across.

Second, the serving economics. If your DeepSeek V4 traffic is long-context and KV bound, halving cache bytes is the lever, and it matters more than another 5 percent of decode speed. My earlier notes on the model's attention design are in DeepSeek V4 GA and the DSA endpoint retirement, and the serving knobs around it in vLLM 0.21 speculative decoding and reasoning budget.

What I would want before trusting it: the same comparison at 64k and 128k context, a retrieval style eval instead of GSM8K, and the concurrency number on a workload where KV really is the limit. Until someone posts that, the honest summary is a promising 2x cache cut with a modest measured benefit. Who has the MI355X time to run the long-context eval?

#vllm#deepseek#nvfp4#rocm