Super IntelligenceDocsHome

Engineering

Reproducibility

A claim about a model is only as strong as another person's ability to check it. Reproducibility is what turns results into evidence, and it is the cheapest way for a new project to earn trust.

State of the field

Tooling. Experiment tracking, declarative configuration and container images pinned by digest are standard.

Disclosure. Model cards and longer system cards accompany major releases, although developers now disclose less about training than they did.

Open weights. Strong open-weight models are available under a range of licences. Some are permissive, such as MIT and Apache 2.0. Others are custom licences with restrictions on use.

Open problems

  • Non-determinism. Differences between hardware and low-level libraries make exact reproduction of large runs impractical.
  • Open weights are not open training. Releasing weights does not release the data or the code that produced them.
  • Cards are not audited. Model and system cards have no standard form and no independent check.

Record every run

FieldPurpose
Git commitThe exact code
Configuration hashThe exact settings
Data manifest hashThe exact data
Random seedsRepeatable sampling
Library versionsRepeatable numerics
HardwareExplains remaining differences

A run record looks like this. The values are illustrative.

{
  "run_id": "2026-10-01-a3f9",
  "git_sha": "4c1e9b7",
  "config_hash": "sha256:9d2f...",
  "data_manifest_hash": "sha256:71ab...",
  "seeds": { "data": 17, "init": 3, "sampling": 101 },
  "libraries": { "torch": "2.7.1", "inspect_ai": "0.3.120" },
  "hardware": "8x accelerator, single node",
  "energy_kwh": 412.6
}

Publish logs, not only scores

With each model card, publish the full evaluation logs. State the licence and the terms of acceptable use precisely.

Write a system card

A system card describes the whole deployed system, not the model alone.

  1. Intended use and uses that are out of scope.
  2. Evaluation results with intervals, and the harness used.
  3. Safety evaluations and the thresholds applied.
  4. Known limitations and failure modes.
  5. Safeguards in place, and what they do not cover.
  6. Changes since the previous version.

Read licences in full

Before depending on an open-weight model, read the licence text itself. Summaries and round-ups are often wrong about restrictions.

Reproduce before building on a result

Rerun a published baseline before comparing against it. If the number cannot be reproduced, report that, and compare against the number that was obtained.