RRSI: your self-improving agent harness is overfitting to its own train set
If you let an agent rewrite its own prompts, tools and memory rules against a fixed set of tasks, the score on those tasks goes up. Obviously. The useful question is what happens on the tasks it never saw, and a paper that was the top Hugging Face daily paper on September 22 (55 upvotes) finally puts numbers on it. It's called RRSI, regularized recursive self-improvement, arXiv 2609.24972, from Google Research. Fair warning: I worked from the summaries and the headline design, not the full paper, so read the mechanism below as my interpretation and not a replication.
Every accepted edit is a parameter
The thing being evolved is the harness (prompt, control flow, tooling, memory), not the weights. A proposer suggests edits, a selector keeps the ones that score better on the train tasks, and you loop. The failure is the boring classic. The harness learns the quirks of your train set. A prompt line like "the result file is always called out.json" is a real gain on your 50 tasks and clutter, or a bug, on everyone else's.
ML people know this in their sleep. Harness people forget it because edits are text, and text feels like engineering instead of fitting. It isn't. Two hundred accepted edits against fifty tasks is a lot of free parameters, and nobody is holding any of them out.
Brakes on both sides of the loop
RRSI puts one brake on the proposer and two on the selector. The proposer gets an annealed edit budget: the number of edits it may make per round shrinks as the run goes on. That's what annealing normally means, though I don't know their exact schedule. Early rounds can be bold, late rounds only get to nudge. The selector gets a critic and a pruner. Going by the names, the critic asks whether a proposed edit is a general fix or a patch for one train task, and the pruner deletes accumulated edits that no longer earn their place.
The reported result against unregularized evolution: +14.1 points in-distribution, +4.7 points across five out-of-distribution benchmarks, and 30% fewer policy tokens. I don't know what baseline the +14.1 is measured from, so check that before quoting it. Note the shape of it, though. Roughly two thirds of the in-distribution gain does not travel (4.7 against 14.1). That gap is the overfitting, and it's the honest number in the paper.
A self-improving harness without a holdout set is just a very slow way to memorize your benchmark.
Where the 30% probably comes from
I suspect a good part of the token saving is plain deletion. Unregularized evolution accretes. Each accepted edit adds a sentence to the prompt, a paragraph to a tool description, another rule in memory, and you pay for all of it on every single call. A pruner that removes dead weight cuts the bill without touching the model.
A made-up fleet to make it concrete: 50M policy tokens a day at a blended $6 per million is $300 a day. Thirty percent off is $90 a day, about $2,700 a month, from removing text. And that's before you count the second effect: a shorter, stabler prefix caches better. (Any edit near the top of the prompt invalidates the cache, which is another argument for a small edit budget late in the run.)
What I would copy into a home-grown loop
You don't need Google's setup for the discipline. Split your tasks three ways before the first round: train, a holdout the loop scores on, and a set that nothing touches until the end. Cap the edits per round and let the cap fall. Keep the harness in git with one commit per accepted edit and a note on which task motivated it, so pruning is just a revert plus a rerun.
rounds: 12
edit_budget: 8 -> 1 # linear decay per round
accept_if: train_gain > 0 and holdout_gain >= 0
every 3 rounds: revert each edit in turn, keep the revert if holdout does not drop
That block is mine, not the paper's. It's the cheapest approximation of the critic and pruner I can think of, and the acceptance rule alone (holdout must not get worse) is the part I'd bet on. I haven't run this loop myself.
This is the same problem I wrote about in verification is the new bottleneck, one level up. The selector is a verifier, and a verifier you can overfit to will be overfit to. If your loop's only signal is the train score, you've built a generator with no reviewer.
The question I can't shake is a smaller one. How many of the lines in my current CLAUDE.md and system prompts would survive a pruner that reverts each one and reruns the holdout? I'd guess fewer than half, and I'd rather find out on a Tuesday than after the bill.