In 1965 the Bletchley Park cryptologist I. J. Good , who had worked alongside Turing, wrote that once a machine can design its successors, an intelligence explosion begins and human intellect is left behind. The idea now travels as RSI (recursive self-improvement — a system improving its own ability to improve) , and it splits into two readings: easy RSI automates R&D and replaces human researchers, hard RSI makes the model itself more compute or data efficient by directly rewriting its weights or architecture. Lilian Weng's July 4 Harness Engineering survey sharpens that split by reviewing about 35 papers and arguing the real lever is not the model but the harness — the evaluation, selection and orchestration scaffolding around the model . Like a good workshop where the bench outlives the craftsman, the harness carried this year's math headlines.
Beijing's Five-Stage Map: The Last AI Built by Humans
Last week 33 researchers led by ByteDance, Tsinghua and the Shanghai AI Laboratory published The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement , grading RSI into five stages. 1) AI merely executes a human-written improvement procedure, 2) it chooses how to improve itself, 3) it decides what new information or experience to acquire, 4) it adapts post-deployment to distribution shift, 5) it persistently refines the very methods used to improve AI itself. The authors stress a filter: fixing a single response is not RSI; improvements must persist beyond one task and be inherited by successors. The competitive stake is clear: automating the labor-heavy life cycle of training and tuning shortens cycles and cuts cost, read in the US-China race as squeezing more performance from scarce hardware .
The Real Engine: AlphaEvolve's Discovery Loop
The summer's math headlines — the Jacobian conjecture, the Navier-Stokes push and progress on the Riemann hypothesis — share a pattern: no model walked alone. Google DeepMind's AlphaEvolve (May 2025) institutionalized it. Gemini Flash scans breadth, Gemini Pro adds depth, while automated evaluators score each program and an evolutionary program database decides which ideas seed the next prompt. The loop is three steps: 1) propose — the agent writes code, 2) measure — the evaluator logs score and crash, 3) select and feed back — the most promising programs shape the next prompt. The payoff spilled outside the lab: data-center efficiency, chip design and even the LLM training process itself improved under the same loop. The surprise is not in the model, but in the harness steering it.
The Bottleneck No One Mentions: Exploration Policy
Inside every loop hides a quiet decision maker: the exploration policy — the rule set that decides where to branch next, what to run in parallel and when to cut a dead line . The video illustrates it with horse matching — if one attempt pairs horses slightly better, should the agent double down or restart? If an attempt crashes before a single match, was the idea or the implementation at fault? Until last week that policy was hard-coded by whoever set up the loop; a fixed strategy cannot learn from accumulated experience and keeps paying for directions that already failed. Like hitting the same side wall in a maze, it re-bills failure. Two walls rise together: policy feedback is delayed and expensive — you must watch a whole discovery run to judge a policy — and the policy space is vast, so most candidates you would need to try are bad ones costing a full rollout to expose.
How Dream RSI Dreams in Three Steps
Seventeen researchers from Google, DeepMind, Maryland and Virginia take aim at that bottleneck in Dream-RSI: Recursive Self-Improvement through Evolving Worlds (arXiv:2609.14858, Sept 14) without touching weights. It works in the orchestration layer and repeats in three steps: 1) Online exploration — a frozen coding agent ( Gemini 3.1 Pro / 3.7 Flash in the paper) solves a real problem while each decision is stored as a discovery tree node carrying code, score and crash flag. 2) Replay simulator — a lightweight simulator that treats the recorded tree as a deterministic world — a finished run is no longer text but a structured tree; a new policy never re-executes anything, it walks the same tree in a different order and every outcome it asks for is already on disk. This simulator is not a learned world model; it is exact where history went, with no prediction. 3) Dreaming policy refinement — a fixed policy-developer agent generates dozens of candidate policies, each replayed at zero execution cost across all historic trees; the scoring formula rewards solution quality, penalizes revealed nodes (real cost) and rewards parallelism, and the highest dream scorer is redeployed online. Every deployment grows a new tree, so the selected policy is the one that wins across a pool of worlds, not the luck of one run.
Why is this cheap? Because the bill is already paid. When a policy's only job is to re-read history in a different order, thousands of candidates cost as much as walking records on disk instead of thousands of full runs. The paper's calling card says the team collects histories online, builds simulators from them, refines the meta-exploration strategy via dreaming and redeploys it. Like memorizing a chess game's move tree and then trying thousands of variants in your head without touching the board, you find the best opening risk-free. The limit comes from the same place: the simulator is exact only where history visited; a branch never visited cannot be dreamed, which is why the loop must keep producing fresh history.
What Changed Across Eight Tasks: Lasso and 162x Fewer Calls
The team ran the same setup twice — fixed versus dreaming — on eight tasks in algorithm design and mathematics . The headline is a lasso solver (sparse regression — the algorithm that picks the simplest accurate model among hundreds of variables, like cooking the best dish with the fewest ingredients) that beats the standard Python machine learning library ; the dreaming policy found it in about 300 tries , the fixed policy needed 550 , and the prior best method took roughly 51,000 . The paper and independent summaries report this as up to 162x fewer search calls . A small prompt detail matters: the agent is explicitly asked to read every past attempt before writing code , to avoid tiny tweaks to the same idea and to promise not to kill processes — a tiny wording change that yields a large strategic shift. For example , on a 500-feature dataset each attempt proposes a new coefficient-pruning program; the dreaming policy discards a pruning family that crashed 200 times in history without ever fielding it again, while the fixed policy still spends 40 more tries inside the same family.
AI commentary
"What matters to me is not the model getting larger but the search-control layer maturing in dreams; when weights stay frozen and strategy learns, the intelligence-explosion rhetoric cools into an engineering ledger."
AI assessment
The video's strongest synthesis gets this right: dreaming the policy while weights stay frozen is not a Good-style intelligence explosion , yet it is a measurable lever in engineering. At its best it mirrors a grad researcher revising the morning strategy against the night's logs — the model stays equally clever, but search waste falls. That steelman reframes Dream RSI from hype to durable improvement at the harness layer .
Limits are crisp. The simulator is exact only where history visited ; a brilliant idea never tried cannot be dreamed. Evaluation also stays tied to hand-written scorers — beating sklearn on lasso is meaningful, yet if the scorer is narrow the policy can game it and miss real generalization. Finally, when the evaluator itself is model-based, as with o3-style reasoners, the cost math shifts; even if dreaming is free, final validation remains expensive.
Provenance splits. Fireship's narration compresses Dream RSI into one video, while arXiv:2609.14858 and ByteDance's five-stage map put different RSI definitions side by side in the same week — one measures orchestration, the other an autonomy ladder. Independent replay needs three checks: a) are code and scorers open for the eight tasks behind the 162x claim, b) does the same dream scoring formula hold across combinatorial versus continuous optimization, c) is the policy tuned pool-wide or to the luck of one tree? Without those, explosion rhetoric is premature.
Practical takeaways diverge. For math and algorithm teams , Dream RSI is immediately usable with a frozen coding agent wherever the evaluator is cheap and deterministic — nightly logs become morning strategy. For product and general-purpose agents with expensive, open-ended or human-feedback evaluators, the dream pool shallows quickly and human-in-the-loop remains the cost driver. Who benefits thus hinges on how cheap and trustworthy your scorer is.
Sources
7 links; 2 of them also cited by 3 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Fireship: Did Google just kickstart the intelligence explosion?
- @arxiv.org https://arxiv.org/abs/2609.14858
Also cited by: The costume on Arena: Gemini 4 traces leaking under the Gemini 3.8 Flash label · Google's Dream RSI: Turning History into a Simulator for Recursive Self-Improvement
- @aimodeling.com https://www.aimodeling.com/en/news/slug/dream-rsi-replay-simulator
Also cited by: Gemini 4 Pro Leaks, GPT-6 Soul in Testing: From Arena to Google Cloud, the Week's AI Shockwave
- @shattered.io https://shattered.io/dream-rsi-recursive-self-improvement-2026/
- @deepmind.google https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/
- @thestar.com.my https://www.thestar.com.my/aseanplus/aseanplus-news/2026/09/16/chinese-researchers-chart-5-stage-path-toward-last-ai-built-by-humans
- @channelnewsasia.com https://www.channelnewsasia.com/east-asia/chinese-researchers-map-out-five-stage-plan-ai-improve-without-human-intervention-6393996
dream rsi · rsi · alpha evolve · google deepmind · bytedance