Can AI improve AI? We read 223 papers to find out.
The literature is fragmented ā long-horizon agents, AI4AI, self-improvement, and RSI are four separate conversations that don't cite each other. So we organized all of it around one question: how far can an AI system reliably carry an improvement process from idea to verified result?
Two findings:
1. AI is now excellent at the *work* of improvement ā planning, coding, experimentation, repair. Humans still set the goals and decide what counts as progress. That asymmetry isn't closing.
2. The composition gap: strong performance on individual components rarely composes into reliable end-to-end improvement. A system can beat every component benchmark and still fail the full loop. Most "self-improving AI" claims live in this gap.
Can AI improve AI? We read 223 papers to find out.
The literature is fragmented ā long-horizon agents, AI4AI, self-improvement, and RSI are four separate conversations that don't cite each other. So we organized all of it around one question: how far can an AI system reliably carry an improvement process from idea to verified result?
Two findings:
1. AI is now excellent at the *work* of improvement ā planning, coding, experimentation, repair. Humans still set the goals and decide what counts as progress. That asymmetry isn't closing.
2. The composition gap: strong performance on individual components rarely composes into reliable end-to-end improvement. A system can beat every component benchmark and still fail the full loop. Most "self-improving AI" claims live in this gap.