- weights change
- model rewrites itself
+ search policy rewrites itself
+ 42% fewer agent callsRecursive self-improvement is the idea that an AI makes itself smarter, then uses the smarter self to do it again, and this week Google published a paper with those two words in the title, in which the model's weights never move. Three numbers to hold onto. Anthropic's CEO says this loop has been running since the summer, which is his reason to slow the whole industry down. The tweet announcing that Google just cracked it collected eight hundred likes in a day. And the paper's own headline result is 317 model calls instead of 550, which is a discount, not a takeoff.
In three minutes: where the phrase came from, what actually loops inside Dream-RSI, and how to read the next paper with RSI on the cover. Start in 1965. I. J. Good, a Bletchley Park statistician who worked next to Turing, writes that an ultraintelligent machine could design even better machines, calls it an intelligence explosion and the last invention man need ever make, provided the machine is docile enough to tell us how to control it, which is the provided-that people stop reading after the comma. For forty years the phrase meant exactly that: a system that edits its own source or weights and comes out smarter every pass. Eliezer Yudkowsky's 2008 essay made it the load-bearing word of AI safety, and nobody had a machine to test it on, which is the ideal condition for a definition.
Then in May 2025 DeepMind shipped AlphaEvolve, a Gemini agent that proposes code, gets it scored, keeps the winners and mutates them again. It multiplied four-by-four complex matrices in 48 steps, beating a record Strassen set in 1969, and it recovers about zero point seven percent of Google's worldwide compute, which at Google's scale is a data centre found under the sofa. Gemini was now helping train Gemini, and the phrase quietly moved from safety blogs into press releases. Dream-RSI bolts a controller on top of that loop. A policy, an actual Python program, decides which branches of the search tree to expand, how many in parallel, and when to give up. Gemini writes the candidate code, and a fixed evaluator scores it. Every run leaves behind a tree of what was tried and what it scored. That tree becomes a replay simulator: a second, equally frozen Gemini rewrites the policy, and the new policy is tested against the recorded outcomes instead of on real GPUs. The paper calls this dreaming, and one real run pays for thousands of dreamed ones at zero execution cost.
The winner goes back online, the pool of recorded worlds grows, repeat. And the prompt in Appendix B.2 tells the self-improving AI, in writing: edit only this one file, and do not solve the scientific task. Now the measurements. On the Lasso solver task, fixed exploration spent 550 Gemini calls to reach three point six seconds, Dream-RSI spent 317 to reach two point nine, and both beat scikit-learn. On KernelBench it reached the same kernel speed with about two point four times fewer generations. And on circle packing it scored two point six three five nine eight three, which is precisely, to six decimals, what AlphaEvolve version two scored last year.
So X read the title: Google just cracked recursive self-improvement. Hacker News read the PDF: calling this RSI seems misleading. And the same week a competitor's CEO said the loop had been running since summer and everyone should slow down. Three sentences, the same two words, three different machines. So, Monday, how to read the next RSI paper. One: ask what object changes, the weights, the code, or the search settings; Dream-RSI changes the third. Two: ask who is frozen; here it's the coder, the judge and the rewriter, all three. Three: ask what the headline number is a unit of. Compute saved is optimisation; a benchmark the model couldn't reach before is the takeoff, and nobody has published that one yet.
Verdict, under the hood: needs review. The loop is real, the savings are measured, and the two words on the cover describe a machine that isn't in the paper.
Verdict: NEEDS REVIEW — Real loop, real savings. The weights never move.
Sources
https://arxiv.org/abs/2609.14858
https://arxiv.org/pdf/2609.14858
https://github.com/zhengkid/Dream-RSI
https://dream-rsi.com
https://news.ycombinator.com/item?id=49726955
https://huggingface.co/papers/2609.14858
https://en.wikipedia.org/wiki/I._J._Good
https://www.lesswrong.com/posts/JBadX7rwdcRFzGuju/recursive-self-improvement
https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/
https://darioamodei.com/post/we-must-pace-the-frontier
https://news.ycombinator.com/item?id=49672510
https://x.com/thesupermannx/status/2100153430849499395
https://x.com/Reezxy23/status/2100348119405707507
And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
YouTube · thedailydiff.dev · forward this to the intern who deployed on Friday.

