Research
Reasoning and autonomy
Two abilities separate a capable assistant from an autonomous researcher: reasoning that can be checked, and reliable work over long tasks. Both have advanced quickly, and both have clear limits today.
Reasoning and test-time compute
State of the field
Chain-of-thought prompting showed that models reason better when they write intermediate steps. From 2024, models have been trained to do this, trading accuracy against the number of thinking tokens.
Search has returned through mathematics. AlphaProof, a reinforcement-learning prover working in the Lean proof language, reached silver-medal standard at the 2024 International Mathematical Olympiad together with AlphaGeometry 2. In 2025 an advanced version of Gemini Deep Think scored 35 of 42 points, gold-medal standard, working in natural language within the contest time and graded officially.
AlphaEvolve pairs language models with evolutionary search and automatic evaluators. It found a way to multiply 4×4 complex matrices with 48 scalar multiplications.
Open problems
- Faithfulness. Whether the visible reasoning reflects the computation that produced the answer.
- Depth. One study reports accuracy collapsing at high problem complexity. Critics attribute the result to output limits and missing tools. The dispute is unresolved.
- Verification outside formal domains. Lean is a ground-truth checker for mathematics. Most of science and engineering has none.
- Cost. Top results often need orders of magnitude more inference compute per problem.
Research directions
- Search over reasoning traces with learned value functions for partial solutions.
- Generator and verifier pairs, using formal provers, executable tests and simulation as checkers.
- Autoformalisation at research level, with libraries that grow from machine-proved lemmas.
- Measures of how faithful and how monitorable a reasoning trace is.
Long-horizon autonomy
State of the field
METR measures the length of task, in human working time, that a model completes at 50% reliability.
| Measure | Value | Source |
|---|---|---|
| Doubling time, 2019–2025 | about 7 months | Kwa et al., 2025 |
| Doubling time, after 2023 | about 131 days (95% interval 107–161) | Time Horizon 1.1 |
| Frontier estimate, spring 2026 | about 16–17 hours (interval 8.5–55 hours) | METR time horizons |
METR cautions that measurements above 16 hours are unreliable with its current task suite. The frontier estimate should be read with that in mind.
Open problems
- Errors compound. Agent success fits a constant failure rate per unit of time, a half-life. Horizons at 80% or 99% reliability are therefore much shorter than at 50%.
- External validity. The tasks are mostly software, self-contained and scored automatically. Horizons differ by domain.
- Benchmarks and the field diverge. In a randomised trial, experienced open-source developers were 19% slower with early-2025 AI tools, while believing they were faster.
- Self-monitoring. Knowing when to stop, recover or ask.
Research directions
- Reliability-centred metrics: horizons at 80, 95 and 99%, failure-rate curves, and benchmarks of recovery from error.
- Task suites lasting weeks, with measured human baselines and underspecified goals.
- Hierarchical planning, checkpointing, self-verification and review by a second agent.
- Field experiments alongside benchmarks.
Recommended practice
- Report reliability at more than one threshold, never the 50% horizon alone.
- Pair every generator with a verifier that the generator cannot edit.
- Measure uplift with controlled trials before claiming productivity gains.