← all posts
// local · local

Ternary Bonsai 27B: how much of the 1.58 bits is RAM saved and how much is quality paid

The number that pushed PrismML's Ternary-Bonsai-2-27B to the top of Hugging Face trending is 1.58, and I'd like to take it apart before it lands in somebody's slide deck. The GGUF is built on Qwen3.8, sat at about 1,275 likes and 1.5 million downloads when the numbers I have were taken (last update September 17), and showed up on OpenRouter on the 18th. I haven't run it. Everything below is arithmetic from the format plus a measurement plan, not a benchmark.

Where 1.58 comes from and what it buys in RAM

Ternary means every weight is one of three values: -1, 0 or +1. Three states carry log2(3) ≈ 1.585 bits. So the ideal footprint for 27 billion weights is 27e9 × 1.585 / 8 ≈ 5.35 GB. The same weights in fp16 are 54 GB. A Q4_K_M quant of a 27B model, at roughly 4.85 bits per weight on average (my memory of llama.cpp's figure, so treat it as approximate), lands around 16.4 GB.

Real files never hit the ideal. The ternary types I know in llama.cpp, TQ1_0 and TQ2_0, spend about 1.69 and 2.06 bits per weight if I remember the block layouts right, which puts 27B at roughly 6 to 7 GB. Which packing Bonsai uses I haven't checked. Embeddings and the output layer are often kept at higher precision too, and the KV cache doesn't shrink at all when the weights do. So "runs in a few GB" is honestly "6 to 8 GB of weights, plus whatever your context costs".

Still, going from 54 GB to about 7 is real. It moves a 27B model from workstation territory to a laptop with 16 GB and some patience.

Bandwidth is the better half of the story

Token-by-token decoding reads every weight once per token, so the ceiling is simple: tokens per second is at most memory bandwidth divided by weight bytes. Take an imaginary machine with 400 GB/s (I picked the number, it's not a spec of anything specific). fp16 gives 400 / 54 ≈ 7 tok/s. Q4_K_M gives 400 / 16.4 ≈ 24. A ternary file at 5.35 GB gives 75, and at 6.95 GB about 58.

Those are ceilings. Attention reads the KV cache, and ternary weights still have to be unpacked and multiplied, which costs compute that a plain 8-bit path doesn't pay. On a fast-memory chip the unpack step can become the limit before the bus does. That is where the kernel work matters.

Ternary shrinks the bytes for certain; whether it also shrinks the wait depends on how cheaply the unpacking runs on your silicon.

Two llama.cpp changes landed around the same time, and I want to be careful not to oversell them. PR #29153 adds vectorised NEON kernels named q8_K_4x4 and q8_K_4x8 for ARM64. The brief I work from says they give a measurable speedup with unchanged gemm results and gives no numbers, so I have none either. My reading of the names is that they concern the 8-bit activation side with an interleaved layout, which mostly helps prefill, and I can't tell you if they touch the ternary path at all. PR #29136 fixes Metal deprecation warnings from the macOS 27 SDK. That is housekeeping so your build stays clean, not a speedup.

The half you can't compute from a spec sheet

Bytes are deterministic. Quality is not. What I don't know is whether Bonsai was trained with ternary weights in the loop or converted after the fact, and that is the whole ballgame. Squeezing a finished 27B model down to three values per weight without retraining would hurt badly. A model trained for it is a different story. The trending count doesn't help me here: 1.5 million downloads measures curiosity, not accuracy.

My expectation, and it's only an expectation, is the usual pattern with aggressive quantisation. Chit-chat and summaries hold up, while long-context recall, code and structured tool calls slip first. If you plan to run agents on it, those are exactly the things to test.

A test that answers the actual question

The comparison that matters is at equal bytes, not equal parameter count. Fitting 27B into 7 GB is impressive only if it beats what else fits into 7 GB, say a 9B to 14B dense model at Q4 or Q5. So I'd run something like this on the same Apple Silicon machine, everything fully offloaded, five repetitions each.

llama-bench -m bonsai-27b-ternary.gguf -p 512 -n 128 -ngl 99 -r 5
llama-bench -m qwen-27b-Q4_K_M.gguf -p 512 -n 128 -ngl 99 -r 5
llama-bench -m dense-12b-Q4_K_M.gguf -p 512 -n 128 -ngl 99 -r 5
llama-perplexity -m <each>.gguf -f wiki.test.raw -c 2048

The first column of results to read is tg128 against the bandwidth ceiling above: how close does each file get? The second is prompt processing (pp512), where ternary has no built-in advantage. Then perplexity as a rough sanity check, though it is a weak proxy across different architectures. Then the part I trust most: a hundred of your own prompts including a few tool-call schemas, scored by a human. And watch resident memory at 32k context, because that's where the small file stops looking small.

If the 27B ternary wins that three-way race at 7 GB, the hype is earned. If it ties the 12B, you got a bigger model for the same cost and nothing else. Which one is it? I'd rather see the numbers than guess.

#local#quantization#llama.cpp#apple-silicon