Super IntelligenceDocsHome

Research

Measuring progress

A field that cannot measure its progress cannot steer it. Public benchmarks now saturate within one to two years of release, and several have known defects. This page lists what the main benchmarks show, how they fail, and how to measure with more care.

The main benchmarks

BenchmarkWhat it measuresKnown weakness
METR time horizonsLength of task completed at 50% reliabilityMostly software tasks; unreliable above 16 hours; few long tasks have measured human baselines
SWE-benchResolving real GitHub issuesContamination and flawed tests reported in the Verified subset
SWE-bench ProLonger software tasks, 1,865 tasks across 41 repositoriesNewer, less studied
FrontierMathResearch-level mathematics, 338 problems in version 2135 problems needed correction; the funder holds exclusive access to a subset
Humanity's Last ExamClosed-ended expert questionsExam scores are not open-ended research
GAIAGeneral assistant tasks; 92% human against 15% GPT-4 at launchSaturating
ARC-AGILearning new tasks from few examplesEarlier versions well represented in training data
RE-BenchResearch engineering against human expertsSmall task set, short budgets

How benchmarks fail

  • Contamination. Public tasks enter training data.
  • Construct validity. A closed-ended exam score is not the same thing as open-ended work.
  • Label and test errors. Graders and reference answers contain mistakes.
  • Scaffold and budget effects. The same model scores very differently under different agent software and token budgets, so leaderboards mix model and harness.
  • Independence. A funder with access to held-out data weakens trust in the result.
  • Ceilings. Once the best systems solve most items, the benchmark stops separating them.

Statistical practice

Most published comparisons report a single number. That is not enough to tell a real difference from noise.

  • Report the number of items, and a standard error or 95% confidence interval.
  • When two models answer the same items, analyse the paired differences.
  • Cluster errors when items are grouped, for example several questions on one passage.
  • Run a power analysis before the evaluation to decide how many items are needed.

These follow Miller, 2024.

Research directions

  • Private, rotating, procedurally generated test sets held by a third party.
  • Reporting normalised for cost and reliability: score per unit cost, variance, and success on repeated attempts.
  • Aggregation across benchmarks with item-response theory.
  • Real-world outcome studies: randomised trials, deployment data and externally verified discoveries.
  1. Publish full evaluation logs, not only scores.
  2. State the harness, tool set, token budget and number of attempts with every agent result.
  3. Keep a private held-out set and embed canary strings to detect leakage.
  4. Never reuse a parent model's results for a quantised or distilled version.

The engineering side of this is covered in Evaluation harness.