The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement · Yi Duan, Fan Wu and 33 co-authors, corresponding author Xuanhe Zhou · arXiv, posted September 10, 2026, revised September 15
A 35-author roadmap defines five levels of recursive self-improvement
The survey builds a five-level autonomy framework, running from a system that executes human-specified updates to one that revises the mechanisms governing its own future improvement, and uses it to classify existing work. The Darwin Gödel Machine, which evolves coding agents and raised its score on a SWE-bench subset from 20 percent to 50 percent, still leaves its archive-maintenance and parent-selection rules outside the system's own control, so the paper credits it with better outputs but not with a self-revised improvement process. Gödel Agent, which rewrites both its task policy and its own improvement logic, ended 14 percent of its 100 recorded MGSM optimization trials below where it started. The authors use that result to argue that persistence alone is not evidence of progress. The paper also cites first-party operational figures: Anthropic reports agentic workloads use roughly four times the tokens of ordinary chat, rising to about fifteen times for multi-agent systems, and OpenAI reports its internal coding-inference compute grew 100-fold over six months of GPT-5.6 development. On Hacker News, in the same thread discussing Dream-RSI, several commenters made the same point this paper makes explicitly: repeated use of "RSI" and "AGI" without a shared definition lets very different claims travel under one banner.
Why it matters: The framework separates systems that produce better outputs from systems that revise the mechanism producing those outputs. Applied to Dream-RSI, that distinction separates a better exploration controller from a system that governs every part of its future improvement.