Super IntelligenceDocsHome

Engineering

Data

Data decides what a model can learn and what it is licensed to know. The fastest-growing category is no longer text. It is environments: tasks in containers with programs that grade the result.

State of the field

Corpora. Pretraining mixes web crawl, licensed collections, code and a growing share of synthetic data such as model-written reasoning traces and tool-use records.

Environments. Reinforcement learning for agents needs tasks that can be attempted many times and graded automatically. Open examples include Reasoning Gym and GEM.

Limits of human text. The stock of high-quality human-written text is finite.

Open problems

  • Provenance and licensing. Litigation and opt-out rules are unsettled, and dataset documentation lags behind.
  • Feedback loops. Training on model output narrows the distribution over time.
  • Leakage. Benchmark items find their way into training corpora.
  • Reward hacking. Agents exploit grader bugs, and environment quality is hard to audit at scale.
  • Poisoning. A small number of planted documents can install a backdoor.

Keep a manifest

Every source gets one row. Every training sample traces back to a row.

FieldExample
SourceName and URL or supplier
LicenceExact licence or contract reference
CollectedDate of crawl or delivery
HashContent hash of the stored snapshot
FiltersDeduplication, quality and safety filters applied

Filter and deduplicate

  • Remove exact and near duplicates.
  • Filter for quality with a classifier, and record its version.
  • Maintain a blocklist of evaluation canary strings, and drop any document containing one.

Build environments like software

  1. Use fixed seeds and sealed containers so that a run can be repeated.
  2. Keep grader tests out of the agent's file system.
  3. Before training, attack the grader with an agent told to cheat.
  4. Version every environment. A changed grader is a new environment.

Guard against poisoning

  • Record where each document came from and when.
  • Scan for backdoor triggers before and after training.
  • Hold back a clean reference set to compare behaviour against.