← all posts
// models · models

Naive-N0.5-Flash: 48 layers, no full attention, and what that does to a 1M-token cache

NaiveAI released Naive-N0.5-Flash this weekend: 309B parameters in a mixture of experts, 15.5B active per token, MIT license, native 1M context, FP8 weights on Hugging Face. It's built on Xiaomi's MiMo-V2.5, and the company describes it as AI-assisted research, meaning AI building frontier AI. The part I care about is the layer stack. Of 48 transformer layers, 39 use sliding-window attention with a 128-token window, and 9 use DeepSeek Sparse Attention with top-2048 selection. There isn't a single full-attention layer in it.

What each layer type actually costs

A sliding-window layer only looks back 128 tokens. Its KV cache stops growing after 128 entries, no matter how long the conversation gets, and its attention compute per token is constant. That's 39 of the 48 layers with a fixed-size cache.

The nine DSA layers are different, and this is where I'd correct the tidy version of the story. In sparse attention of the DeepSeek kind, each token still gets scored against the earlier ones by a cheap indexer, and then attends to only the top 2048. Compute for the attention proper is bounded. But the keys and values of the whole history still have to live somewhere, because any old token can be picked. So those nine layers keep a cache that grows linearly with context, and the indexer pass is still linear per token. Prefill is cheaper than full attention, not free.

The cache arithmetic, with an assumption

Here is the back-of-envelope. If all 48 layers were full attention, cache size would scale as 48 × n, with n the context length. For this model it scales as roughly 9 × n + 39 × 128. At n = 1,000,000 that's about 9.005 million versus 48 million units, so around 5.3 times smaller.

The assumption hiding in that number is that every layer stores the same amount per token. The material I worked from doesn't give head counts, head dimensions, or whether the DSA layers use a compressed latent cache the way DeepSeek's own do, so I can't turn units into gigabytes. If you want that conversion for your hardware, the KV cache math piece has the formula. Plug in the real config from the model card, not my ratio.

Thirty-nine windows of 128 is not a memory

There's a modeling question the architecture raises, and it's the one I'd want answered before trusting the 1M label. A 128-token window is small. Information can hop further by stacking layers, so in principle the receptive field of the window layers is about 39 × 128, close to 5,000 tokens. Anything beyond that has to be found by one of nine sparse layers, and each of them picks 2048 candidates out of up to a million.

That's a tight budget for retrieval. Needle-in-a-haystack tests are the easy version of the problem. What I'd want are multi-hop questions where the needle is a fact you only recognize as relevant after reading something else, spread across 800K tokens. Context windows are a lie was about this gap between advertised and usable length, and native 1M on paper doesn't close it.

Removing full attention makes long context affordable. Whether it stays useful is a separate claim that needs its own evidence.

What the release doesn't tell me

The speed figures come from NaiveAI's own inference stack, NaiveRT: 50 tokens per second per user in the standard mode and up to 2,000 in an ultrafast mode. I don't know the hardware, batch size, context length or whether the ultrafast mode leans on speculative decoding, so I wouldn't compare those numbers with anything. The benchmark list is agentic coding and research (SWE-Bench Pro, Terminal-Bench 2.1, MLE-bench-30, PaperBench), and I haven't seen independent reproductions. I also haven't run the weights myself, and at 309B even FP8 needs a serious multi-GPU box, so the Apple Silicon question is open.

The first thing I'd test is quality as a function of length on the same task, at 32K, 128K, 512K and 1M. If the curve stays flat, this design deserves the copying it's going to get. If it sags after 200K, the 1M in the name is a cache-size claim, not a usefulness claim.

#models#architecture#local