Super IntelligenceDocsHome

Engineering

Reference architecture

This page proposes how a system built toward advanced AI should be laid out. It is a design, not a description of running software. Each layer links to the page that covers it.

Design principles

  1. Measure before building. The evaluation harness comes first, because nothing else can be judged without it.
  2. Assume the agent can be compromised. Limits are enforced by the system around the model, not by the model's good behaviour.
  3. Everything is reproducible. A result that cannot be rerun from recorded inputs is not a result.
  4. State lives outside the model. Progress, memory and decisions are stored where they can be inspected.
  5. People approve what cannot be undone.

Layers

LayerResponsibilityPage
ComputeAccelerators, scheduling, checkpointing, energy accountingCompute and training
DataSources, provenance, filtering, training environmentsData
TrainingPretraining, fine-tuning, reinforcement learning, distillationCompute and training
ServingInference engine, caching, batching, cost trackingInference
Agent runtimeThe loop, tools, memory, sub-agentsAgents
EvaluationTasks, scoring, statistics, regression gatesEvaluation harness
SafetySandbox, credentials, monitoring, approval, incident responseSafety engineering
RecordsRun tracking, configuration, model and system cardsReproducibility

How a request flows

user or scheduler
      │  task
      ▼
┌──────────────┐   tool call    ┌───────────────────┐
│ agent runtime│ ─────────────▶ │ policy gate       │  allow / deny / ask a person
│ (model loop) │ ◀───────────── │ (outside the model)│
└──────┬───────┘   result       └─────────┬─────────┘
       │                                  │ allowed call
       │ every step                       ▼
       ▼                          ┌───────────────┐
┌──────────────┐                  │ sandbox       │  no ambient secrets,
│ audit log    │ ◀─────────────── │ (microVM)     │  egress allow-list
│ (append-only)│                  └───────────────┘
└──────┬───────┘
       ▼
  monitors and reviewers

The model never calls a tool directly. Every call passes a policy gate that the model cannot edit, and every step is written to a log that the model cannot change.

Release gates

A change moves forward only when it passes each gate in order.

GatePasses when
1. RegressionCapability, safety behaviour and calibration suites show no significant drop
2. Dangerous capabilityResults stay below the thresholds set in advance
3. ControlRed-team exercises show the protocol holds for the new abilities
4. ReviewA named person signs off, with the evaluation logs attached

Gates 2 and 3 follow Evaluations and Control and oversight.

What is deliberately left open

  • Train or build on open weights. The architecture supports both. The choice is listed as undecided in Project status.
  • Scale. Nothing here assumes frontier-scale compute. The same structure applies to small experiments.