Research
Automated AI research
The idea that separates superintelligence from every earlier technology is recursion: a system that helps design its successor. This page separates what has been measured from what is forecast.
The idea
In 1965 I. J. Good observed that a machine able to surpass people at every intellectual activity could also surpass them at designing machines, and so improve itself. He called the result an intelligence explosion. The open question is the speed of the loop, usually called takeoff speed.
What has been measured
RE-Bench compares AI agents with human experts on research-engineering tasks under a time budget.
| Time budget | Result |
|---|---|
| 2 hours | Agents scored about four times higher than human experts |
| 8 hours | Humans narrowly ahead |
| 32 hours | Humans scored about twice the best agent |
Agents are fast and strong on short tasks. People still lead where a task needs sustained direction.
Field evidence is more cautious. In a randomised trial, experienced developers were 19% slower with early-2025 tools and believed they were faster. Perceived and measured uplift can diverge.
What is forecast
The scenario AI 2027 described superhuman coders in 2027. By January 2026 its authors gave later medians, December 2030 and January 2035, and stressed that 2027 had never been a confident forecast.
In a 2023 survey of 2,778 AI researchers, the median respondent gave a 20% chance that superintelligence follows human-level machine intelligence within two years, and 80% within thirty.
Companies have published their own estimates of how much of their research is automated. These are self-measured and not externally audited, and this documentation does not rely on them.
Open problems
- Taste. Whether choosing ideas and experiments can be automated, or only implementation.
- Compute bottlenecks. If experiments, not labour, limit progress, automated researchers add less acceleration. This is the centre of the takeoff-speed disagreement.
- Diminishing returns. Whether each doubling of capability needs more than double the research input.
- Measurement. Self-reported automation shares have no common definition.
Research directions
- Open successors to RE-Bench with multi-day budgets and scoring for novelty.
- Independent, repeated uplift trials inside research workflows.
- Empirical estimates of the returns to software research and of how far labour substitutes for compute.
- A standard taxonomy of automation levels, audited by third parties.
Why this matters for safety
A loop that speeds itself up shortens the time available to notice and correct errors. For that reason the safety work in this documentation is designed to run before and during capability work, not after it. See Control and oversight.
Recommended practice
- Treat any system that contributes to its own training pipeline as a higher-risk system, with stricter evaluation and approval.
- Log every automated change to training code, data or configuration, with a human owner.
- Publish the method used to measure research automation along with the number.