Engineering
Data
Data decides what a model can learn and what it is licensed to know. The fastest-growing category is no longer text. It is environments: tasks in containers with programs that grade the result.
State of the field
Corpora. Pretraining mixes web crawl, licensed collections, code and a growing share of synthetic data such as model-written reasoning traces and tool-use records.
Environments. Reinforcement learning for agents needs tasks that can be attempted many times and graded automatically. Open examples include Reasoning Gym and GEM.
Limits of human text. The stock of high-quality human-written text is finite.
Open problems
- Provenance and licensing. Litigation and opt-out rules are unsettled, and dataset documentation lags behind.
- Feedback loops. Training on model output narrows the distribution over time.
- Leakage. Benchmark items find their way into training corpora.
- Reward hacking. Agents exploit grader bugs, and environment quality is hard to audit at scale.
- Poisoning. A small number of planted documents can install a backdoor.
Recommended practice
Keep a manifest
Every source gets one row. Every training sample traces back to a row.
| Field | Example |
|---|---|
| Source | Name and URL or supplier |
| Licence | Exact licence or contract reference |
| Collected | Date of crawl or delivery |
| Hash | Content hash of the stored snapshot |
| Filters | Deduplication, quality and safety filters applied |
Filter and deduplicate
- Remove exact and near duplicates.
- Filter for quality with a classifier, and record its version.
- Maintain a blocklist of evaluation canary strings, and drop any document containing one.
Build environments like software
- Use fixed seeds and sealed containers so that a run can be repeated.
- Keep grader tests out of the agent's file system.
- Before training, attack the grader with an agent told to cheat.
- Version every environment. A changed grader is a new environment.
Guard against poisoning
- Record where each document came from and when.
- Scan for backdoor triggers before and after training.
- Hold back a clean reference set to compare behaviour against.