Super IntelligenceDocsHome

Safety

Interpretability

Mechanistic interpretability tries to read what a model computes from its internal activity. It is the only approach that looks inside the system instead of at its behaviour. It is also immature, and its limits need stating as clearly as its results.

State of the field

Superposition. Networks represent more concepts than they have neurons by overlapping them. Individual neurons are therefore hard to read.

Dictionary learning. Sparse autoencoders separate overlapping activity into more interpretable features, and the method scales to production models.

Circuit tracing. Attribution graphs trace how features combine to produce one specific output.

A survey of the remaining gaps is in Sharkey et al., 2025.

Open problems

  • Reconstruction error. Feature dictionaries do not capture all of a model's activity, and there is no ground truth for the right set of features.
  • Local explanations. An attribution graph explains one prompt, and it explains a simplified replacement model, not the original.
  • Decodable is not used. A probe shows that information can be read from activations, not that the model relies on it.
  • No completeness guarantee. Failing to find a deception feature is not evidence that none exists.
  • Evasion. Tools have not been tested against models trained to defeat them.

What it can and cannot certify

Can do todayCannot do today
Generate hypotheses about how a behaviour arisesCertify the absence of misaligned goals
Find concepts nobody thought to look forCertify the absence of planted behaviour
Debug a specific failureGive a global guarantee about a frontier model
Supply signals for runtime monitoringReplace behavioural evaluation
  1. Use it as one line of evidence. Combine interpretability findings with behavioural evaluations in any alignment assessment.
  2. Compare against simple baselines. Test any feature-based method against linear probes and random baselines before relying on it.
  3. Run auditing games. One team plants a hidden behaviour, another tries to find it with interpretability tools. Report the detection rate.
  4. Deploy probes as monitors with measured error. A probe used at runtime needs a known false-negative rate.