Research
Measuring progress
A field that cannot measure its progress cannot steer it. Public benchmarks now saturate within one to two years of release, and several have known defects. This page lists what the main benchmarks show, how they fail, and how to measure with more care.
The main benchmarks
| Benchmark | What it measures | Known weakness |
|---|---|---|
| METR time horizons | Length of task completed at 50% reliability | Mostly software tasks; unreliable above 16 hours; few long tasks have measured human baselines |
| SWE-bench | Resolving real GitHub issues | Contamination and flawed tests reported in the Verified subset |
| SWE-bench Pro | Longer software tasks, 1,865 tasks across 41 repositories | Newer, less studied |
| FrontierMath | Research-level mathematics, 338 problems in version 2 | 135 problems needed correction; the funder holds exclusive access to a subset |
| Humanity's Last Exam | Closed-ended expert questions | Exam scores are not open-ended research |
| GAIA | General assistant tasks; 92% human against 15% GPT-4 at launch | Saturating |
| ARC-AGI | Learning new tasks from few examples | Earlier versions well represented in training data |
| RE-Bench | Research engineering against human experts | Small task set, short budgets |
How benchmarks fail
- Contamination. Public tasks enter training data.
- Construct validity. A closed-ended exam score is not the same thing as open-ended work.
- Label and test errors. Graders and reference answers contain mistakes.
- Scaffold and budget effects. The same model scores very differently under different agent software and token budgets, so leaderboards mix model and harness.
- Independence. A funder with access to held-out data weakens trust in the result.
- Ceilings. Once the best systems solve most items, the benchmark stops separating them.
Statistical practice
Most published comparisons report a single number. That is not enough to tell a real difference from noise.
- Report the number of items, and a standard error or 95% confidence interval.
- When two models answer the same items, analyse the paired differences.
- Cluster errors when items are grouped, for example several questions on one passage.
- Run a power analysis before the evaluation to decide how many items are needed.
These follow Miller, 2024.
Research directions
- Private, rotating, procedurally generated test sets held by a third party.
- Reporting normalised for cost and reliability: score per unit cost, variance, and success on repeated attempts.
- Aggregation across benchmarks with item-response theory.
- Real-world outcome studies: randomised trials, deployment data and externally verified discoveries.
Recommended practice
- Publish full evaluation logs, not only scores.
- State the harness, tool set, token budget and number of attempts with every agent result.
- Keep a private held-out set and embed canary strings to detect leakage.
- Never reuse a parent model's results for a quantised or distilled version.
The engineering side of this is covered in Evaluation harness.