Super IntelligenceDocsHome

Engineering

Compute and training

Frontier training has moved from clusters of tens of megawatts to campuses measured in gigawatts. This page covers the infrastructure and the training methods in common use, and the practices that keep large runs correct.

State of the field

Scale. Epoch AI projected five sites crossing about one gigawatt of facility power during 2026.

Energy. Data centres used about 415 TWh in 2024, and the International Energy Agency projects about 945 TWh by 2030. The agency reports that data-centre demand rose 17% in 2025.

Software. The stack is stable. Large models are split across devices with data, tensor and pipeline parallelism. Training uses 16-bit mixed precision and writes sharded checkpoints asynchronously.

Training methods

StagePurpose
PretrainingLearn general structure from large corpora
Supervised fine-tuningTeach instruction following and formats
Preference-based reinforcement learningShape behaviour from human or model judgements
Reinforcement learning with verifiable rewardsImprove reasoning on tasks a program can grade, such as code and mathematics
DistillationProduce smaller, cheaper models from larger ones

Low-rank adaptation is the default for cheap experiments. Quantisation to 8 or 4 bits is standard for deployment.

Open problems

  • Power is the constraint. Grid connection now limits delivery as much as chips do.
  • Hardware fails during every large run. With hundreds of thousands of accelerators, time between failures is shorter than the run. Silent data corruption is a real source of error.
  • Disclosure is falling. Developers have stopped publishing training compute, duration and dataset size for their largest systems, which weakens outside analysis.
  • Verifiable rewards over-optimise. A model trained against a grader learns the grader. Transfer to work that cannot be graded is unclear.
  • Regression between stages. Each stage can undo behaviour learned in the one before.

Treat failure as normal

  • Checkpoint asynchronously to object storage and verify checksums.
  • Test restoring from a checkpoint on a schedule, not only when something breaks.
  • Log loss and gradient norms per shard to catch silent corruption.

Use mixed precision by default

Keep full precision for numerically sensitive reductions.

import torch

with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
    loss = model(batch).loss
loss.backward()

Gate every stage

  • Run a fixed regression suite after each stage: capability, safety behaviour and calibration.
  • Mix reward sources: programmatic, rubric-graded and human spot checks.
  • Read a sample of the highest-reward trajectories by hand to find exploits.
  • Evaluate quantised and distilled models as separate models.

Account for energy

Record for every run the energy used, the facility's power usage effectiveness and the grid region, beside the compute.