llama.cpp --cpu-mtp: what a gigabyte of VRAM costs in draft latency
The llama.cpp batch of 29 September (builds b11236 to b11242) has one flag I care about: --cpu-mtp. It moves the MTP drafters onto the CPU and, according to the digest I'm working from, frees about 1 GB of VRAM. That's the whole claim. I haven't run it, and the source is an aggregated infra digest, not the release notes themselves, so the first job before quoting any of this is to read the PRs.
The same batch also has multimodal input for /v1/embeddings, a Vulkan speedup of about 6.3 % on an RTX 3090, GCC 15 fixes and the first step of a migration to a batch_ext API.
What you save and what it pays for
Multi-token prediction means a small drafter guesses the next few tokens and the main model verifies them in one pass. I wrote about a similar setup on Apple hardware in the Ollama MTP piece. What --cpu-mtp changes is where the drafter's weights and compute live. Instead of sitting in VRAM next to the main model, they sit in system RAM and run on CPU cores.
What does 1 GB buy on a discrete GPU? Depends on the model, so here is one illustrative shape: 32 layers, 8 KV heads, head dimension 128, fp16 cache. KV bytes per token are 2 x 32 x 8 x 128 x 2 = 131,072, which is 128 KiB. One gibibyte is then 8,192 tokens of extra context. Or one more block of layers moved off the CPU. On a 24 GB card already running a tight quant, that's a meaningful margin. On a 48 GB one, it's a rounding error.
On Apple Silicon, I doubt it buys anything. Unified memory means the drafter's weights occupy the same pool whether a CPU or a GPU touches them. The brief frames the freed gigabyte as a win for M-series machines, and I don't see where it comes from, unless a GPU wired-memory limit is the binding constraint. I'm not sure about that one and won't pretend to be.
The draft step gets slower, by how much?
The cost side is latency. A CPU has a fraction of the memory bandwidth of a GPU, and the drafter must finish before the main model can verify. Here is a toy model, every input invented for illustration. The expected tokens accepted per step with draft length k and per-token acceptance rate a is (1 - a^(k+1)) / (1 - a). With a = 0.7 and k = 3 that is (1 - 0.2401) / 0.3, about 2.53 tokens.
Suppose the main model needs 25 ms per verification step, so 40 tokens per second without speculation. A GPU drafter at 1.5 ms per draft token makes the step 25 + 4.5 = 29.5 ms, and 2.53 tokens in 29.5 ms is about 86 tokens per second. A CPU drafter at 6 ms per token makes it 25 + 18 = 43 ms, which is about 59 tokens per second.
So in this toy, the CPU drafter still beats no drafter (59 against 40), but it gives up roughly a third of the speedup compared with GPU drafting (2.15x down to 1.47x). Change the inputs and the picture changes: a faster CPU path, or a drafter that overlaps with verification, closes the gap. A drafter that needs to copy hidden states from the GPU on every step would widen it. I don't know which of those applies here, and the PR will say.
Every gigabyte you hand back comes out of the draft phase's latency budget, so the honest question is what the gigabyte is worth to you.
How I plan to measure it
No numbers from me yet, so here is the plan instead, which you can steal. Fixed model and quant, fixed prompt set with both short and long prompts, same build, same thread count. Run A is my usual MTP setup, run B adds --cpu-mtp. For each: decode tokens per second, draft acceptance rate, time to first token, and VRAM used, read with nvidia-smi --query-gpu=memory.used --format=csv before and during a long generation. I'd also watch CPU load, because the drafter's threads compete with whatever else the server is doing, prompt processing included.
The result worth reporting is the pair: VRAM saved divided by decode speed lost. If it's 1 GB for 5 % it's a bargain. If it's 1 GB for 30 %, you'd do better dropping a quant level instead. For background on how much memory tuning of this kind is worth at all, see the GPU offload math piece.
The other items in the batch
The /v1/embeddings change lets the endpoint accept structured content arrays, meaning image, audio and video parts, not only text. Useful if you want a single OpenAI-style call for multimodal retrieval. Whether it works for you depends on whether your model actually has an embedding path for those modalities, and I haven't tested any.
The Vulkan figure, +6.3 % on an RTX 3090 with 4096-token batches, is a single number from a single setup. Big batches suggest prefill, so don't expect decoding to move. It is mainly interesting if CUDA isn't an option for you, and 6.3 % is within the range where I'd repeat the run five times before believing it.
Unsloth also reports roughly 4.5x faster image and video generation on Apple Silicon. That's their claim, and I'm leaving it there until someone measures where the speedup comes from.
When I do run --cpu-mtp, the first thing I'll check is whether the drafter needs the GPU's hidden state on each step, because that one detail decides whether the flag is a bargain or a trap.