Super IntelligenceDocsHome

Safety

Control and oversight

Alignment tries to make a system want the right things. Control asks a different question: can the system be used safely even if it wants the wrong things? Oversight asks how people can supervise work they cannot fully check. Both are needed, and both have limits that grow with capability.

AI control

State of the field

The control approach assumes the most capable model may be working against its operators and designs protocols that stay safe anyway.

RoleMeaning
Untrusted modelThe strong model whose intentions are not verified
Trusted modelA weaker model believed unable to scheme
Human auditA small budget of expert review
Red teamReplaces the untrusted model with an attack strategy
Blue teamDesigns the protocol that must stay safe

Common protocols are trusted monitoring, untrusted monitoring with measures against collusion, trusted editing, and deferring suspicious actions to the trusted model.

For agents that take many steps, resampling suspicious actions reduced attack success from 58% to 7% at a 5% cost to usefulness.

Open problems

  • It is a bridge, not a destination. Control does not extend to systems far more capable than the trusted monitor and the human red team.
  • Red-team adequacy. Whether human-designed attacks are a conservative stand-in for a capable adversary.
  • Collusion. Hidden coordination between copies of an untrusted model, and attacks spread thinly across a long deployment.
  • A widening gap. The distance between trusted and untrusted models grows each generation.
  • Attention. Human audit budgets do not grow with deployment.

Scalable oversight

State of the field

  • Debate. Two models argue and a weaker judge decides. Debate between more persuasive models helped weaker judges reach more accurate answers. Other work found the benefit modest and dependent on the task.
  • Weak-to-strong generalisation. A strong model supervised by a weaker one partly recovers its full capability.
  • Recursive reward modelling. Largely conceptual to date.

Open problems

  • Debate assumes honest strategies win. Obscure arguments and exploiting the judge are unsolved.
  • Weak-to-strong results may not hold when the strong model resists.
  • Using AI to do alignment research means trusting the researcher's output, which is circular when alignment is the question.
  1. Treat internal agents as insiders. Least privilege, sandboxing, monitoring of every action, and human approval for high-impact actions.
  2. Evaluate control before granting new abilities. Run red-team and blue-team exercises, and report the trade-off between safety and usefulness.
  3. Log everything centrally. Record all agent actions and audit them with trusted models.
  4. Resample or defer. When an action looks suspicious, sample again or hand it to the trusted model.
  5. Break hidden channels. Paraphrase or redact messages between untrusted components.
  6. Keep automated safety research under the same controls.

The implementation is described in Safety engineering.