Engineering
Safety engineering
This page turns the safety research into system requirements. The starting assumption is uncomfortable and useful: the agent may be hijacked or may be working against its operators, and the system must stay safe anyway.
State of the field
Isolation. Code run by agents is confined in containers hardened with gVisor, or in Firecracker microVMs, with outbound network access denied by default.
Prompt injection. It is treated as unsolved at the level of the model. Adaptive attacks defeat published defences. The field has moved to system designs that separate control flow from untrusted data.
Weights. RAND's security levels give a common scale for protecting model weights.
Open problems
- No robust model-level defence against instructions hidden in content.
- Approval fatigue. Requests for human approval lose value when they are frequent.
- Review at scale. Nobody can read every transcript of a twelve-hour run.
- Insider and supply-chain threats to weights.
Requirements
1. Sandbox every execution
| Control | Requirement |
|---|---|
| Isolation | MicroVM or hardened container per task |
| Network | Deny by default, with an allow-list of destinations |
| File system | Only the task's working directory is writable |
| Lifetime | Destroyed when the task ends |
2. Scope credentials
- Issue short-lived credentials per task, limited to what the task needs.
- Keep no ambient secrets in the sandbox.
- Revoke on completion or on any alert.
3. Put policy outside the model
Every tool call passes a gate the model cannot edit. The gate allows, denies or asks a person.
# policy.yaml — evaluated outside the model on every tool call
defaults:
network: deny
filesystem: workspace-only
rules:
- tool: shell
allow: true
monitor: true
- tool: http_request
allow_hosts: [api.internal.example, pypi.org]
- tool: deploy
require_approval: human
- tool: send_message
when: session.read_untrusted_content
require_approval: human
The last rule applies the three-way rule from Security: a session that has read untrusted content may not communicate externally without a person's approval.
4. Require approval for what cannot be undone
Deployment, payment, deletion and external messages need a person. Keep the list short so that each request is read.
5. Log everything, immutably
Record every tool call with its inputs and outputs in an append-only store. Review with a trusted model, and send a sample to people.
6. Monitor in layers
| Layer | Catches |
|---|---|
| Input and output classifiers | Known harmful content |
| Trusted-model monitor | Suspicious actions in context |
| Activation probes | Internal signals, with a measured false-negative rate |
| Human audit | A sample, plus everything flagged |
7. Prepare for incidents
A written runbook covers four steps: stop the agent, revoke its credentials, preserve the transcripts, and review afterwards. Rehearse it.
8. Protect weights
Encrypt at rest. Restrict and log access. Require two people for any export path. Monitor outbound data volume.