Context Language Models: let the agent edit its own context file
The idea in the Context Language Models paper is almost embarrassingly simple: instead of a harness deciding when to compact your agent's history, you put the context in a file and let the model edit it with Bash. The paper (arXiv 2609.37725, University of Washington with Meta Superintelligence Labs, dated September 30, 2026) is more interesting for the serving trick it needs than for the idea itself.
Context as a file the model can rewrite
A normal LM appends. The paper writes it as c(t+1) = c(t) concatenated with f(c(t)). A CLM instead makes the next context an arbitrary function of the previous one. In practice the context is mirrored into a file, the model can append or rewrite it with shell commands, and every edit is synced back into the live context for the next turn. Several agents can each own a context file, which is how the authors handle subagents and agent swarms.
What I like is what they compare against. Existing harnesses compact at a fixed threshold, or give the model a narrow compaction tool. The authors argue, citing the bitter lesson, that unrestricted edits beat human-designed policies. The paper reports models inventing trackers for multi-agent orchestration, a separate chat role for internal notes, and reusable context-management functions. Those are qualitative examples, and I would want to see how often they appear before getting excited.
The numbers, with the baselines attached
The accuracy story is modest; the compute story is where this paper earns its keep.
Applied zero-shot to Qwen3.6-27B and GPT5.6-Sol, CLM reaches 11.4% higher accuracy than the strongest baseline on BrowseComp-Plus with 21.5% fewer prefix-reuse FLOPs. On TerminalBench 2.1 it matches the best baseline's accuracy with 29.5% fewer FLOPs. On a 10-task subset of 12-hour EdgeBench it scores 5% higher with 59% fewer FLOPs than a Codex-style summarization harness, and on a 24-hour six-repository swarm task it gets 65% greater speedup at the same compute.
Note the subset. Ten tasks is a small sample, and the paper says so itself. Note also that "FLOPs" here is a defined quantity (prefix-reuse FLOPs, covering decoding, prefilling and re-prefilling under standard serving), not a wall-clock or a bill.
The weak spot is small models. Qwen3.5-9B with CLM starts about six points below the summary harness, because a 9B model is not good at managing its own context. Online RL fixes much of it: CLM goes from 28.8% to 42.5% on BrowseComp-Plus, while the summary harness goes from 34.7% to 42.1%. At the end they are level on accuracy, but CLM spends 1.34 PFLOPs per question against 2.19. So the accuracy edge is gone and the cost edge stays.
Why edits hurt your KV cache
Here is the part that matters if you run inference. Prefix caching only helps when the new prompt shares a prefix with the old one. If the model replaces a block B in the middle of [A B C] with B', standard serving matches A and then re-prefills B' and all of C. The longer C is, the worse it gets.
Take an illustrative case with numbers I chose: A is 10K tokens, B is 20K, B' is 2K, C is 60K. Standard prefix reuse re-prefills 62K tokens. Suffix Cache Reuse (SCR), the paper's fix, re-prefills only the 2K of B'. Whatever the real distribution of edits looks like, the structure of the saving is plain: it scales with the size of the untouched tail.
SCR diffs the new prompt against the session's previous prompt, finds surviving spans, and relocates up to K of them (K is 6 in the main text), largest first. For each, it reuses the cached keys and values, re-rotates the rotary position encodings to the new positions, and splices them in after B'. The relocated entries sit in session-private slots, so the shared radix tree never holds a moved entry, and if allocation fails the server falls back to normal re-prefilling.
The catch is stated in the paper: those reused states were computed under the old context, so SCR approximates re-prefilling rather than reproducing it. The authors report matched task performance, and on BrowseComp-Plus with Qwen3.6-27B it runs at 65.0% of standard SGLang's prefix-reuse FLOPs, which is the headline 35% saving. I would still want to test it on my own workloads before trusting an approximation in a long agent run. It is a patch to SGLang, with a separate path for hybrid models that mix full and linear attention layers (the linear layers snapshot a recurrent state instead).
Where this leaves context management
Summary compaction is the thing CLM is meant to replace, and the paper notes that summaries can lose or hallucinate information on tests like Needle Retention and Sudoku Sketchpad. If you have ever watched an agent forget a constraint after a compaction, that sounds familiar (see also context budgets in Claude Code and compacting tool results). I have not run CLM or SCR myself, and the code is listed at github.com/facebookresearch/context-language-models.
For now the practical read is about serving, not prompting. Any harness that lets a model rewrite the middle of its context will pay the re-prefill bill, and SCR shows how large that bill can be. My own question is whether hosted APIs, which I cannot patch, will ever expose anything like it.