Safety
Security
A safe model that can be stolen, poisoned or hijacked is not safe. Security for advanced AI covers three surfaces: the model's inputs, its training data and its weights.
State of the field
Jailbreaks. No general solution exists. Classifier-based defences raise the cost of attack substantially. Adaptive attacks have bypassed twelve published defences with success above 90%.
Prompt injection. An agent that reads untrusted content can be given instructions through it. The more robust defences enforce policy outside the model. CaMeL separates control flow from untrusted data and solved 77% of benchmark tasks with provable security, against 84% with no defence.
Data poisoning. About 250 poisoned documents were enough to plant a backdoor, regardless of model size.
Weight security. RAND defines five security levels for model weights, from protection against amateurs to protection against the most capable state operations. RAND judges the highest level not achievable at present.
Open problems
- Static tests overstate robustness. A defence tested against fixed attacks tells little about an attacker who adapts.
- Any content is an attack surface. Agents with tools and private data can be reached through web pages, files and messages.
- Provenance. The integrity of web-scale training data cannot be checked by hand.
- Upper security levels. They need hardware, facility and personnel measures that no developer has demonstrated in public.
- Insiders. The threat includes the AI systems themselves.
The three-way rule for agents
An agent session is dangerous when it combines all three of these without supervision:
- Exposure to untrusted input.
- Access to private data.
- The ability to communicate externally or change state.
Remove any one and the worst outcomes become much harder to reach.
Recommended practice
- Test with adaptive red teams. Combine human and automated attackers, and run a bug bounty.
- Isolate by capability. Design agents so that data flow is controlled by the system, and do not rely on the model to resist injection.
- Layer defences. Input and output classifiers, asynchronous monitoring, and rapid response.
- Benchmark against the RAND levels. Reduce the number of people and systems that can reach the weights, require more than one person to authorise export, and limit outbound data volume.
- Curate training data. Filter it, record its provenance, and scan for backdoors.