← all posts
// local · local

Colibri streams MoE experts from disk: the bandwidth math behind it

Colibri showed up on GitHub trending this week, written by JustVugg: a minimal inference runtime in plain C, no dependencies, that runs frontier-class mixture-of-experts models by streaming the expert parameters from disk on demand instead of holding them in high-bandwidth memory. I haven't built it or run a single token through it, so everything below is arithmetic on assumed numbers, not a benchmark. I'll say where the assumptions bite.

The idea is less crazy than it sounds, because of what a MoE actually is.

Why MoE is the one architecture where this can work

A dense model touches every weight for every token. Streaming that from an SSD would be absurd. A mixture-of-experts model has a large pool of expert blocks, and a router picks only a few of them per token per layer. The total parameter count sets how much storage you need. The active count sets how much data you have to move per token. Those two numbers can differ by an order of magnitude or more, and Colibri bets on exactly that gap: keep the always-needed parts (attention, router, shared layers) close, and fetch experts when the router asks for them.

So the question is never “does it fit in RAM”. It is “how many bytes per token do I have to pull across the slowest link in the machine, and how fast is that link”.

Capacity is a storage problem, speed is a bandwidth problem, and streaming experts only trades the first for the second.

The arithmetic, with made-up but plausible numbers

Decode is bandwidth-bound: tokens per second is roughly link bandwidth divided by bytes touched per token. Take a hypothetical model where the routed experts for one token add up to 10 GB at your quantization. That figure is invented for illustration, so plug in your own model.

link               bandwidth     10 GB per token   tokens/s (upper bound)
PCIe 4 NVMe        ~7 GB/s       1.43 s            ~0.7
PCIe 5 NVMe        ~12 GB/s      0.83 s            ~1.2
unified memory     ~400 GB/s     0.025 s           ~40

The 7 GB/s is what a good PCIe 4.0 drive is rated for on sequential reads, and real expert fetches are not perfectly sequential, so treat it as a ceiling. The 400 GB/s is Max-class Apple-silicon territory; I went through the newer numbers in the M6 and M5 Ultra bandwidth sizing piece. The gap between the first and last row is roughly 57x. No amount of clever C closes that.

What narrows it is locality. Routers are not uniform; some experts get picked far more often than others, and the OS page cache (or Colibri's own cache, I don't know which it uses) will keep the hot ones in RAM. If 80 % of expert reads hit cache, disk traffic drops from 10 GB to 2 GB per token, which is about 0.29 s on the PCIe 4 drive, or 3.5 tokens per second. Usable for a patient human. Not pleasant. And the 80 % is a number I pulled from the air: hit rate depends on the model, the prompt and how diverse your workload is, and I have no measurement for any of it.

Where the trade makes sense

The disk is slow per token, but it is enormous and cheap per gigabyte. A 2 TB NVMe costs a rounding error next to 2 TB of anything with HBM or even unified memory attached. So the honest use cases are the ones where latency per token doesn't matter and capacity does.

Batch and offline work is the obvious one. When a batch of tokens goes through a layer together, each expert you read serves every token in the batch that routed to it. Bytes per token falls as batch size rises, and the disk link gets amortized. Summarizing a thousand documents overnight is a very different proposition from chatting.

Prefill behaves similarly. A long prompt has so many tokens that nearly every expert gets used at least once, so you read the whole pool once per prompt and spread the cost across thousands of tokens. Decode is the painful phase, which is also the phase people judge speed by. Claims about running big models on ordinary hardware usually quote whichever phase flatters them, so ask which one.

The third case is plain curiosity: you want to poke at a model that you could never otherwise load. I like that one, and it's a legitimate reason.

If you're already thinking about spilling experts out of the fast tier, the same logic drives the H100 plus 2 TB offload experiment, just with system RAM as the slow tier instead of flash.

The claim I can't confirm

The pitch that this runs frontier models on everyday hardware is the part I'd hold at arm's length. Technically it may be true: the model loads, tokens come out. But “runs” hides a range from 0.5 tokens per second to something you'd tolerate, and the repository's own description doesn't tell me which models, which quantization, which drive, or what decode speed. Until somebody posts prefill and decode numbers separately, on named hardware, with cache hit rates, I'd file it as promising and unmeasured.

There's also flash wear to think about. Reading doesn't wear an SSD the way writing does, so that's less scary than it sounds, but sustained random reads will still heat a laptop drive and throttle it. Another thing I haven't tested.

What I'd actually want from the maintainer: a table of tokens per second against cache size, same model, same prompt set. That single chart would tell you whether your machine sits on the useful side of the curve.

#local#hardware#models#analysis