Super IntelligenceDocsHome

Engineering

Safety engineering

This page turns the safety research into system requirements. The starting assumption is uncomfortable and useful: the agent may be hijacked or may be working against its operators, and the system must stay safe anyway.

State of the field

Isolation. Code run by agents is confined in containers hardened with gVisor, or in Firecracker microVMs, with outbound network access denied by default.

Prompt injection. It is treated as unsolved at the level of the model. Adaptive attacks defeat published defences. The field has moved to system designs that separate control flow from untrusted data.

Weights. RAND's security levels give a common scale for protecting model weights.

Open problems

  • No robust model-level defence against instructions hidden in content.
  • Approval fatigue. Requests for human approval lose value when they are frequent.
  • Review at scale. Nobody can read every transcript of a twelve-hour run.
  • Insider and supply-chain threats to weights.

Requirements

1. Sandbox every execution

ControlRequirement
IsolationMicroVM or hardened container per task
NetworkDeny by default, with an allow-list of destinations
File systemOnly the task's working directory is writable
LifetimeDestroyed when the task ends

2. Scope credentials

  • Issue short-lived credentials per task, limited to what the task needs.
  • Keep no ambient secrets in the sandbox.
  • Revoke on completion or on any alert.

3. Put policy outside the model

Every tool call passes a gate the model cannot edit. The gate allows, denies or asks a person.

# policy.yaml — evaluated outside the model on every tool call
defaults:
  network: deny
  filesystem: workspace-only

rules:
  - tool: shell
    allow: true
    monitor: true

  - tool: http_request
    allow_hosts: [api.internal.example, pypi.org]

  - tool: deploy
    require_approval: human

  - tool: send_message
    when: session.read_untrusted_content
    require_approval: human

The last rule applies the three-way rule from Security: a session that has read untrusted content may not communicate externally without a person's approval.

4. Require approval for what cannot be undone

Deployment, payment, deletion and external messages need a person. Keep the list short so that each request is read.

5. Log everything, immutably

Record every tool call with its inputs and outputs in an append-only store. Review with a trusted model, and send a sample to people.

6. Monitor in layers

LayerCatches
Input and output classifiersKnown harmful content
Trusted-model monitorSuspicious actions in context
Activation probesInternal signals, with a measured false-negative rate
Human auditA sample, plus everything flagged

7. Prepare for incidents

A written runbook covers four steps: stop the agent, revoke its credentials, preserve the transcripts, and review afterwards. Rehearse it.

8. Protect weights

Encrypt at rest. Restrict and log access. Require two people for any export path. Monitor outbound data volume.