Super IntelligenceDocsHome

Research

Learning and world models

Today's deployed models have fixed weights, learn little from experience after training, and learn about the physical world mostly from text and video. This page covers three linked gaps: memory, models of the world, and learning efficiency.

Memory and continual learning

State of the field

Persistence today comes from long context windows, retrieval and external notes managed by agent software. Research prototypes aim at learning during use:

  • Titans adds a neural long-term memory that is updated at inference time.
  • SEAL has a model write its own fine-tuning data and update instructions.

Neither is established at frontier scale.

Open problems

  • Catastrophic forgetting. New learning overwrites old when weights are updated in sequence.
  • Context is not learning. Retrieval returns facts, not skills, and performance degrades over very long contexts.
  • Safe learning from deployment. Privacy, poisoning, and drift in values and capability. A model that keeps changing must be re-evaluated.
  • No standard benchmark for learning on the job over weeks.

Research directions

  • Memory on several timescales: fast episodic stores with slow consolidation into weights.
  • Isolated parameter updates, such as adapters, with guarantees against forgetting.
  • Memory managed by the agent, with learned policies for writing, retrieving and forgetting.
  • Assurance methods for models that are not stationary.

World models and embodiment

State of the field

The Joint Embedding Predictive Architecture predicts in a learned latent space instead of predicting pixels or tokens, and plans with the resulting model. V-JEPA 2 was pretrained on more than one million hours of video. With under 62 hours of unlabelled robot video it supports zero-shot pick-and-place planning.

Vision-language-action models now show long household manipulation tasks in homes they have not seen.

Open problems

  • Whether models trained on text and video acquire causal, physical models or surface statistics.
  • Robot data is scarce. No internet-scale dataset of actions exists.
  • Generated worlds lose consistency over long horizons.
  • Hierarchical planning in latent space remains largely unsolved.

Research directions

  • Head-to-head comparison of latent predictive and generative world models on planning tasks.
  • World models as training simulators and as evaluators of robot policies.
  • Benchmarks of intuitive physics and counterfactual prediction.

Sample efficiency and generalisation

State of the field

ARC-AGI tests the ability to learn a new task from a few examples. The results show both progress and the remaining gap.

BenchmarkResultSource
ARC-AGI-2, best open competition entry, 202524.03%ARC Prize 2025
ARC-AGI-2, best verified commercial systems, 202537.6% at $2.20 per task; 54% at $30 per taskARC Prize 2025
ARC-AGI-3 at launch, March 2026Humans 100%, frontier AI 0.51%ARC-AGI-3

ARC-AGI-3 is interactive: no instructions and no stated goal. It measures exploration as well as inference.

Open problems

  • People learn new tasks from a handful of examples. Models need far more data or test-time compute.
  • Models are brittle to surface changes in a problem.
  • Efficiency is under-reported. Score per unit of cost matters as much as score.

Research directions

  • Program synthesis guided by deep learning, and refinement loops.
  • Training and adaptation at test time.
  • Interactive benchmarks that score exploration and action efficiency.
  • Private, procedurally generated test sets that can be renewed.
  1. Report cost per task beside every score.
  2. Re-evaluate a model after any update that changes its weights or its memory.
  3. Keep test sets private and renew them on a schedule.