← all posts
// hardware · apple

512 GB on a desk: a sizing table for what the M5 Ultra Mac Studio can hold

Mac Studio goes on sale today (September 22): M5 Max from $2,499, M5 Ultra from $5,499, with up to 512 GB of unified memory. The catch is that the 512 GB configuration does not ship until the end of October, so for now this is a planning exercise. They check whether the model fits and stop there. Fitting is the easy half.

The bandwidth and compute story is already in M6 and M5 Ultra: 1.2 TB/s and what actually fits, and prefill versus decode is in M5 Neural Accelerators. This is the lookup table I wish I had when someone asks me whether a particular model runs on the big Studio.

Assumptions, stated before the numbers

Weights take parameters times bits divided by eight. A 70B model at 4-bit is 70 × 4 / 8 = 35 GB. Real quantised files are a bit bigger, because scales and biases cost extra bits per group (call it 4.5 effective bits for a typical 4-bit format, which puts that 70B at roughly 39 GB). I ignore that in the table and keep the clean arithmetic, so add about ten percent in your head.

Usable memory is not 512 GB. I assume 384 GB, three quarters, on the theory that macOS limits how much the GPU may wire by default. I have not checked the default on an M5 Ultra, and as far as I know it can be raised with a sysctl, so treat 384 as the conservative case. The decode ceiling divides the 1.2 TB/s figure from my earlier piece by the bytes read per token. It is a ceiling. I have no measured tokens per second for this machine, and you should distrust anyone who has them before the 512 GB units exist.

The table

Model (hypothetical shapes)4-bit weights8-bit weightsFits in 384 GB?Decode ceiling at 4-bit
27B dense13.5 GB27 GBbothabout 89 tok/s
70B dense35 GB70 GBbothabout 34 tok/s
123B dense61.5 GB123 GBbothabout 19 tok/s
405B dense202.5 GB405 GB4-bit onlyabout 6 tok/s
235B MoE, 22B active117.5 GB235 GBbothabout 109 tok/s
1T MoE, 32B active500 GB1,000 GBneithern/a

The 235B row is the shape of Qwen3-235B-A22B, and the 1T row is invented for illustration. Everything else is just arithmetic on the parameter count.

Where bandwidth gets you instead of capacity

Look at the 405B row. At 8-bit it needs 405 GB, which is under 512 and over my 384. Even if you raise the wired limit enough to load it, each token streams the full 405 GB, and 1,200 / 405 is about 3 tokens per second. It fits and it is useless for chat. Capacity answered the wrong question.

MoE flips this. The 235B model holds 117.5 GB at 4-bit but reads only its active experts per token, about 11 GB, which is why its ceiling is higher than the 70B dense model. You pay for capacity once and for bandwidth on active parameters only. That is the real argument for a 512 GB box: sparse models with enormous total size and small active size. A 2.8T-parameter model such as Kimi K3 is 1,400 GB at 4-bit and still misses, and even at ternary (1.58 bits, 2.8T × 1.58 / 8, about 553 GB) it does not fit in 512. A ternary 27B, for what it is worth, is about 5 GB.

The second bandwidth trap is the KV cache, which people forget to put in the sizing at all. Take a hypothetical model with 80 layers, 8 KV heads, head dimension 128 and an fp16 cache. That is 2 × 80 × 8 × 128 × 2 bytes, or 327,680 bytes per token. At 128k context it is about 43 GB. It is capacity you must budget, and it is also traffic: every decode step reads it. Put that beside the 70B 4-bit weights (35 GB) and the per-token read is 78 GB, which drops the ceiling from 34 to roughly 15 tokens per second. Long context eats decode speed in a way the headline spec sheet never shows.

A model that fits in memory has only passed the first exam; the second is how many gigabytes it has to read for every single token.

Clustering and the cheaper boxes

Thunderbolt 5 clustering lets you pool memory across machines, and I can see why it tempts. But TB5 is 80 Gb/s, roughly 10 GB/s, against 1,200 GB/s inside one Ultra. That is about 120 times less. Splitting a model across boxes works for capacity, and the traffic that crosses the cable has to be small (activations between pipeline stages, not weights), so expect it to serve a model that does not fit, not to make one that fits any faster.

Smaller machines are the other obvious path. The M5 Max Studio starts at $2,499 and the new Mac mini M6 at $899. I have no memory ceiling for the Max from my sources, so look that up before you plan around it. My rough rule stays the same: pick the memory for the largest model you will actually keep loaded, pick the bandwidth for how long you are willing to wait, and run the division above before you click buy.

#apple#m5-ultra#quantization#local