Engineering
Reference architecture
This page proposes how a system built toward advanced AI should be laid out. It is a design, not a description of running software. Each layer links to the page that covers it.
Design principles
- Measure before building. The evaluation harness comes first, because nothing else can be judged without it.
- Assume the agent can be compromised. Limits are enforced by the system around the model, not by the model's good behaviour.
- Everything is reproducible. A result that cannot be rerun from recorded inputs is not a result.
- State lives outside the model. Progress, memory and decisions are stored where they can be inspected.
- People approve what cannot be undone.
Layers
| Layer | Responsibility | Page |
|---|---|---|
| Compute | Accelerators, scheduling, checkpointing, energy accounting | Compute and training |
| Data | Sources, provenance, filtering, training environments | Data |
| Training | Pretraining, fine-tuning, reinforcement learning, distillation | Compute and training |
| Serving | Inference engine, caching, batching, cost tracking | Inference |
| Agent runtime | The loop, tools, memory, sub-agents | Agents |
| Evaluation | Tasks, scoring, statistics, regression gates | Evaluation harness |
| Safety | Sandbox, credentials, monitoring, approval, incident response | Safety engineering |
| Records | Run tracking, configuration, model and system cards | Reproducibility |
How a request flows
user or scheduler
│ task
▼
┌──────────────┐ tool call ┌───────────────────┐
│ agent runtime│ ─────────────▶ │ policy gate │ allow / deny / ask a person
│ (model loop) │ ◀───────────── │ (outside the model)│
└──────┬───────┘ result └─────────┬─────────┘
│ │ allowed call
│ every step ▼
▼ ┌───────────────┐
┌──────────────┐ │ sandbox │ no ambient secrets,
│ audit log │ ◀─────────────── │ (microVM) │ egress allow-list
│ (append-only)│ └───────────────┘
└──────┬───────┘
▼
monitors and reviewers
The model never calls a tool directly. Every call passes a policy gate that the model cannot edit, and every step is written to a log that the model cannot change.
Release gates
A change moves forward only when it passes each gate in order.
| Gate | Passes when |
|---|---|
| 1. Regression | Capability, safety behaviour and calibration suites show no significant drop |
| 2. Dangerous capability | Results stay below the thresholds set in advance |
| 3. Control | Red-team exercises show the protocol holds for the new abilities |
| 4. Review | A named person signs off, with the evaluation logs attached |
Gates 2 and 3 follow Evaluations and Control and oversight.
What is deliberately left open
- Train or build on open weights. The architecture supports both. The choice is listed as undecided in Project status.
- Scale. Nothing here assumes frontier-scale compute. The same structure applies to small experiments.