Research
The capability frontier
Progress toward general and superhuman systems rests on three inputs: compute, data and algorithms. This page summarises what is known about each, where the trend may bend, and which research would settle the open questions.
State of the field
Scaling laws. Language-model loss falls as a power law in parameters, data and compute. For a fixed compute budget, parameters and training tokens should grow in roughly equal proportion.
Compute. Frontier training compute has grown about four to five times per year. Epoch AI judges that training runs near 2e29 FLOP are feasible by 2030 given constraints on power, chips, data and latency.
Algorithms. The compute needed for a fixed level of language-model performance halved about every eight months over 2012–2023, with a 95% interval of 5 to 14 months.
Energy. Data centres used about 415 TWh of electricity in 2024, about 1.5% of the global total, and the International Energy Agency projects about 945 TWh by 2030.
Open problems
- Is algorithmic progress independent of scale? Much of the measured efficiency gain appears to depend on scale and is dominated by the move from LSTMs to Transformers. Progress at small scale is far slower than headline figures suggest.
- Loss is not capability. Scaling laws predict loss. The mapping from loss to reasoning, agency or research ability is empirical and poorly understood.
- Data is finite. High-quality human text is limited, and the value of synthetic data is contested.
- Attribution is missing. Gains since 2024 mix pretraining scale, reinforcement learning after pretraining, test-time compute and scaffolding. No accepted method separates them.
- Physical limits. Power and grid connection, not chips alone, now limit how fast clusters are delivered.
Research directions
- Scaling laws for reasoning. Measure capability against thinking tokens and reinforcement-learning compute, as earlier work did for parameters and data.
- Controlled ablations at several scales. Test each algorithmic change at more than one scale, so that an "effective compute" measure survives the scale-dependence critique.
- Data efficiency. Curricula, active data selection and a theory of when synthetic data helps.
- Public accounting. Reproducible compute-to-capability accounting that covers inference as well as training.
The skeptical case
Four positions deserve a fair statement. A research agenda that cannot answer them is incomplete.
| Position | Claim |
|---|---|
| Yann LeCun | Autoregressive language models lack grounded world models, persistent memory and hierarchical planning. Human-level AI needs predictive models learned from sensory data. |
| Gary Marcus | Pure scaling gives diminishing returns. Reliability and out-of-distribution failures persist, and recent gains come from adding symbolic tools and search. |
| Melanie Mitchell | Benchmark accuracy overstates understanding. Altered versions of familiar tasks suggest much apparent reasoning is approximate retrieval. |
| François Chollet | Skill is not intelligence. Intelligence is the efficiency of acquiring new skills, which scaling memorised skills does not produce. |
Two questions follow from these. Is rising benchmark performance evidence of general reasoning, or of wider coverage of the test distribution? And what observation would change each side's mind? Few pre-registered predictions exist. Adversarial collaborations with tests agreed in advance are the most direct remedy.
Recommended practice
- Report results with the compute spent on training and on inference.
- Test every claimed improvement at two or more scales before treating it as general.
- State in advance what result would count against the hypothesis being tested.