Engineering
Reproducibility
A claim about a model is only as strong as another person's ability to check it. Reproducibility is what turns results into evidence, and it is the cheapest way for a new project to earn trust.
State of the field
Tooling. Experiment tracking, declarative configuration and container images pinned by digest are standard.
Disclosure. Model cards and longer system cards accompany major releases, although developers now disclose less about training than they did.
Open weights. Strong open-weight models are available under a range of licences. Some are permissive, such as MIT and Apache 2.0. Others are custom licences with restrictions on use.
Open problems
- Non-determinism. Differences between hardware and low-level libraries make exact reproduction of large runs impractical.
- Open weights are not open training. Releasing weights does not release the data or the code that produced them.
- Cards are not audited. Model and system cards have no standard form and no independent check.
Recommended practice
Record every run
| Field | Purpose |
|---|---|
| Git commit | The exact code |
| Configuration hash | The exact settings |
| Data manifest hash | The exact data |
| Random seeds | Repeatable sampling |
| Library versions | Repeatable numerics |
| Hardware | Explains remaining differences |
A run record looks like this. The values are illustrative.
{
"run_id": "2026-10-01-a3f9",
"git_sha": "4c1e9b7",
"config_hash": "sha256:9d2f...",
"data_manifest_hash": "sha256:71ab...",
"seeds": { "data": 17, "init": 3, "sampling": 101 },
"libraries": { "torch": "2.7.1", "inspect_ai": "0.3.120" },
"hardware": "8x accelerator, single node",
"energy_kwh": 412.6
}
Publish logs, not only scores
With each model card, publish the full evaluation logs. State the licence and the terms of acceptable use precisely.
Write a system card
A system card describes the whole deployed system, not the model alone.
- Intended use and uses that are out of scope.
- Evaluation results with intervals, and the harness used.
- Safety evaluations and the thresholds applied.
- Known limitations and failure modes.
- Safeguards in place, and what they do not cover.
- Changes since the previous version.
Read licences in full
Before depending on an open-weight model, read the licence text itself. Summaries and round-ups are often wrong about restrictions.
Reproduce before building on a result
Rerun a published baseline before comparing against it. If the number cannot be reproduced, report that, and compare against the number that was obtained.