Engineering
Evaluation harness
The harness is the software that turns a question about a model into a number with an error bar. It is built first and run on every change. The scientific side is in Measuring progress and Evaluations.
State of the field
Frameworks. The common open framework is Inspect from the UK AI Security Institute. An evaluation is a task made of three parts: a dataset, a solver and a scorer. Tasks can run in sandboxes.
Statistics. Good practice is set out in Miller, 2024: standard errors, clustered errors for grouped items, paired comparisons and power analysis.
A minimal task
This follows the Inspect documentation.
from inspect_ai import Task, task
from inspect_ai.dataset import FieldSpec, hf_dataset
from inspect_ai.scorer import model_graded_qa
from inspect_ai.solver import generate
@task
def simpleqa():
return Task(
dataset=hf_dataset(
"codelion/SimpleQA-Verified",
split="train",
sample_fields=FieldSpec(input="problem", target="answer"),
),
solver=generate(),
scorer=model_graded_qa(),
)
inspect eval simpleqa.py --model anthropic/claude-opus-5-5
Open problems
- Saturation and contamination. Public benchmarks stop separating models, and their items leak into training data.
- Human baselines. They are costly and inconsistent. Times for long tasks are often estimated, not measured.
- Model graders. Scoring with a model adds bias and drift.
- Awareness. Models may behave differently when they detect a test.
Recommended practice
Report uncertainty
| Report | Why |
|---|---|
| Number of items | Lets the reader judge precision |
| Standard error or 95% interval | Separates real differences from noise |
| Paired differences | More power when two models answer the same items |
| Clustering by task family | Avoids overstating precision on grouped items |
A comparison without an interval is an anecdote.
Control contamination
- Keep a private held-out set and rotate it.
- Embed canary strings, and search training data for them.
- Check for word-for-word reproduction of reference solutions.
Pin everything
Record the harness version, prompts, sampling parameters and tool versions in the log. Store full transcripts, not only scores.
Run evaluations as continuous integration
| Trigger | Suite |
|---|---|
| Every commit | A small smoke suite, minutes |
| Every release candidate | The full suite |
| Every new ability or tool | Dangerous-capability and control evaluations |
Validate the grader
When a model scores answers, measure its agreement with human graders on a sample, and repeat when the grading model changes.