Super IntelligenceDocsHome

Engineering

Evaluation harness

The harness is the software that turns a question about a model into a number with an error bar. It is built first and run on every change. The scientific side is in Measuring progress and Evaluations.

State of the field

Frameworks. The common open framework is Inspect from the UK AI Security Institute. An evaluation is a task made of three parts: a dataset, a solver and a scorer. Tasks can run in sandboxes.

Statistics. Good practice is set out in Miller, 2024: standard errors, clustered errors for grouped items, paired comparisons and power analysis.

A minimal task

This follows the Inspect documentation.

from inspect_ai import Task, task
from inspect_ai.dataset import FieldSpec, hf_dataset
from inspect_ai.scorer import model_graded_qa
from inspect_ai.solver import generate


@task
def simpleqa():
    return Task(
        dataset=hf_dataset(
            "codelion/SimpleQA-Verified",
            split="train",
            sample_fields=FieldSpec(input="problem", target="answer"),
        ),
        solver=generate(),
        scorer=model_graded_qa(),
    )
inspect eval simpleqa.py --model anthropic/claude-opus-5-5

Open problems

  • Saturation and contamination. Public benchmarks stop separating models, and their items leak into training data.
  • Human baselines. They are costly and inconsistent. Times for long tasks are often estimated, not measured.
  • Model graders. Scoring with a model adds bias and drift.
  • Awareness. Models may behave differently when they detect a test.

Report uncertainty

ReportWhy
Number of itemsLets the reader judge precision
Standard error or 95% intervalSeparates real differences from noise
Paired differencesMore power when two models answer the same items
Clustering by task familyAvoids overstating precision on grouped items

A comparison without an interval is an anecdote.

Control contamination

  • Keep a private held-out set and rotate it.
  • Embed canary strings, and search training data for them.
  • Check for word-for-word reproduction of reference solutions.

Pin everything

Record the harness version, prompts, sampling parameters and tool versions in the log. Store full transcripts, not only scores.

Run evaluations as continuous integration

TriggerSuite
Every commitA small smoke suite, minutes
Every release candidateThe full suite
Every new ability or toolDangerous-capability and control evaluations

Validate the grader

When a model scores answers, measure its agreement with human graders on a sample, and repeat when the grading model changes.