Safety
Control and oversight
Alignment tries to make a system want the right things. Control asks a different question: can the system be used safely even if it wants the wrong things? Oversight asks how people can supervise work they cannot fully check. Both are needed, and both have limits that grow with capability.
AI control
State of the field
The control approach assumes the most capable model may be working against its operators and designs protocols that stay safe anyway.
| Role | Meaning |
|---|---|
| Untrusted model | The strong model whose intentions are not verified |
| Trusted model | A weaker model believed unable to scheme |
| Human audit | A small budget of expert review |
| Red team | Replaces the untrusted model with an attack strategy |
| Blue team | Designs the protocol that must stay safe |
Common protocols are trusted monitoring, untrusted monitoring with measures against collusion, trusted editing, and deferring suspicious actions to the trusted model.
For agents that take many steps, resampling suspicious actions reduced attack success from 58% to 7% at a 5% cost to usefulness.
Open problems
- It is a bridge, not a destination. Control does not extend to systems far more capable than the trusted monitor and the human red team.
- Red-team adequacy. Whether human-designed attacks are a conservative stand-in for a capable adversary.
- Collusion. Hidden coordination between copies of an untrusted model, and attacks spread thinly across a long deployment.
- A widening gap. The distance between trusted and untrusted models grows each generation.
- Attention. Human audit budgets do not grow with deployment.
Scalable oversight
State of the field
- Debate. Two models argue and a weaker judge decides. Debate between more persuasive models helped weaker judges reach more accurate answers. Other work found the benefit modest and dependent on the task.
- Weak-to-strong generalisation. A strong model supervised by a weaker one partly recovers its full capability.
- Recursive reward modelling. Largely conceptual to date.
Open problems
- Debate assumes honest strategies win. Obscure arguments and exploiting the judge are unsolved.
- Weak-to-strong results may not hold when the strong model resists.
- Using AI to do alignment research means trusting the researcher's output, which is circular when alignment is the question.
Recommended practice
- Treat internal agents as insiders. Least privilege, sandboxing, monitoring of every action, and human approval for high-impact actions.
- Evaluate control before granting new abilities. Run red-team and blue-team exercises, and report the trade-off between safety and usefulness.
- Log everything centrally. Record all agent actions and audit them with trusted models.
- Resample or defer. When an action looks suspicious, sample again or hand it to the trusted model.
- Break hidden channels. Paraphrase or redact messages between untrusted components.
- Keep automated safety research under the same controls.
The implementation is described in Safety engineering.