Engineering
Compute and training
Frontier training has moved from clusters of tens of megawatts to campuses measured in gigawatts. This page covers the infrastructure and the training methods in common use, and the practices that keep large runs correct.
State of the field
Scale. Epoch AI projected five sites crossing about one gigawatt of facility power during 2026.
Energy. Data centres used about 415 TWh in 2024, and the International Energy Agency projects about 945 TWh by 2030. The agency reports that data-centre demand rose 17% in 2025.
Software. The stack is stable. Large models are split across devices with data, tensor and pipeline parallelism. Training uses 16-bit mixed precision and writes sharded checkpoints asynchronously.
Training methods
| Stage | Purpose |
|---|---|
| Pretraining | Learn general structure from large corpora |
| Supervised fine-tuning | Teach instruction following and formats |
| Preference-based reinforcement learning | Shape behaviour from human or model judgements |
| Reinforcement learning with verifiable rewards | Improve reasoning on tasks a program can grade, such as code and mathematics |
| Distillation | Produce smaller, cheaper models from larger ones |
Low-rank adaptation is the default for cheap experiments. Quantisation to 8 or 4 bits is standard for deployment.
Open problems
- Power is the constraint. Grid connection now limits delivery as much as chips do.
- Hardware fails during every large run. With hundreds of thousands of accelerators, time between failures is shorter than the run. Silent data corruption is a real source of error.
- Disclosure is falling. Developers have stopped publishing training compute, duration and dataset size for their largest systems, which weakens outside analysis.
- Verifiable rewards over-optimise. A model trained against a grader learns the grader. Transfer to work that cannot be graded is unclear.
- Regression between stages. Each stage can undo behaviour learned in the one before.
Recommended practice
Treat failure as normal
- Checkpoint asynchronously to object storage and verify checksums.
- Test restoring from a checkpoint on a schedule, not only when something breaks.
- Log loss and gradient norms per shard to catch silent corruption.
Use mixed precision by default
Keep full precision for numerically sensitive reductions.
import torch
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
loss = model(batch).loss
loss.backward()
Gate every stage
- Run a fixed regression suite after each stage: capability, safety behaviour and calibration.
- Mix reward sources: programmatic, rubric-graded and human spot checks.
- Read a sample of the highest-reward trajectories by hand to find exploits.
- Evaluate quantised and distilled models as separate models.
Account for energy
Record for every run the energy used, the facility's power usage effectiveness and the grid region, beside the compute.