Safety
Interpretability
Mechanistic interpretability tries to read what a model computes from its internal activity. It is the only approach that looks inside the system instead of at its behaviour. It is also immature, and its limits need stating as clearly as its results.
State of the field
Superposition. Networks represent more concepts than they have neurons by overlapping them. Individual neurons are therefore hard to read.
Dictionary learning. Sparse autoencoders separate overlapping activity into more interpretable features, and the method scales to production models.
Circuit tracing. Attribution graphs trace how features combine to produce one specific output.
A survey of the remaining gaps is in Sharkey et al., 2025.
Open problems
- Reconstruction error. Feature dictionaries do not capture all of a model's activity, and there is no ground truth for the right set of features.
- Local explanations. An attribution graph explains one prompt, and it explains a simplified replacement model, not the original.
- Decodable is not used. A probe shows that information can be read from activations, not that the model relies on it.
- No completeness guarantee. Failing to find a deception feature is not evidence that none exists.
- Evasion. Tools have not been tested against models trained to defeat them.
What it can and cannot certify
| Can do today | Cannot do today |
|---|---|
| Generate hypotheses about how a behaviour arises | Certify the absence of misaligned goals |
| Find concepts nobody thought to look for | Certify the absence of planted behaviour |
| Debug a specific failure | Give a global guarantee about a frontier model |
| Supply signals for runtime monitoring | Replace behavioural evaluation |
Recommended practice
- Use it as one line of evidence. Combine interpretability findings with behavioural evaluations in any alignment assessment.
- Compare against simple baselines. Test any feature-based method against linear probes and random baselines before relying on it.
- Run auditing games. One team plants a hidden behaviour, another tries to find it with interpretability tools. Report the detection rate.
- Deploy probes as monitors with measured error. A probe used at runtime needs a known false-negative rate.