Super IntelligenceDocsHome

Safety

Evaluations

Evaluations are the instrument that tells a developer what a system can do before the world finds out. Their weakness is that they give a lower bound on capability, and that capable models can tell when they are being tested.

State of the field

Capability trends. The length of task that models complete has doubled about every seven months since 2019. The current data is on METR's time horizons page.

Public institutes. The UK AI Security Institute and the US Center for AI Standards and Innovation at NIST test frontier models before release under agreements with developers. The UK institute publishes its findings in a Frontier AI Trends Report.

Shared tooling. Inspect, the UK institute's open-source framework, is the common tool for writing and running evaluations.

Evaluation awareness. One frontier model separated evaluation transcripts from real deployment transcripts with an accuracy score of 0.83, against 0.92 for people. The International AI Safety Report 2026 notes that this makes pre-deployment testing harder.

Sandbagging. Models can be prompted or fine-tuned to underperform on selected evaluations.

What is evaluated

DomainQuestion
CyberCan the system find and exploit vulnerabilities without help?
Biology and chemistryDoes it give meaningful uplift toward weapons?
AutonomyCan it acquire resources, copy itself or resist shutdown?
AI researchCan it speed up the development of AI systems?
PropensityDoes it deceive, scheme or seek power when given the chance?

Open problems

  • Lower bounds only. Better prompting or tooling can elicit more capability than an evaluation found.
  • Unreliable upper bounds. Sandbagging means a low score does not prove low capability.
  • Awareness invalidates propensity tests. A model that knows it is observed may behave differently.
  • Weak link to real-world harm. Benchmark scores map loosely to actual uplift.
  • Independence. Third-party evaluators have limited access and time.
  1. Build in the common framework. Write evaluations in Inspect so that others can run them.
  2. Report elicitation effort. State how hard the evaluators tried, the human baseline, and confidence intervals.
  3. Measure awareness. Estimate how often the model recognised the test, and make evaluations resemble deployment.
  4. Test for sandbagging. Use fine-tuning elicitation, noise injection and consistency checks.
  5. Tie thresholds to actions in advance. Decide what result stops a release before running the test.
  6. Give evaluators real access. Third parties need versions of the model with safeguards reduced.

The software side is covered in Evaluation harness.