Super IntelligenceDocsHome

Reference

Glossary

Terms as they are used in this documentation.

Capability

Narrow AI. A system that performs one kind of task well. Everything in use today.

General AI (AGI). A system that matches people across most mental work. No agreed test exists.

Superintelligence (ASI). A system that clearly exceeds the best people in almost every field. Bostrom distinguishes speed, collective and quality forms (Bostrom, 2014).

Intelligence explosion. The feedback loop in which a system improves the systems that follow it. Proposed by I. J. Good in 1965.

Takeoff speed. The time from human-level AI to superintelligence.

Time horizon. The length of task, in human working time, that a model completes at a given reliability. Usually quoted at 50%.

Test-time compute. Computation spent while answering, such as extended reasoning, as opposed to during training.

Scaling laws. Empirical relations between a model's loss and its parameters, data and compute.

Systems

Agent. A model running in a loop with tools, acting toward a goal.

Harness or scaffold. The software around a model that supplies tools, memory and control flow.

Sub-agent. An agent started by another agent to handle part of a task in its own context.

Context window. The text a model can attend to at once.

Compaction. Summarising earlier context so that work can continue past the window limit.

Retrieval. Fetching stored text into context when it is needed.

Model Context Protocol. An open interface between models and tools or data.

Open weights. A model whose parameters are published. This does not imply open training data or code.

Distillation. Training a smaller model to reproduce a larger one's behaviour.

Quantisation. Storing weights at lower numerical precision to save memory and time.

Safety

Alignment. Making a system pursue what people intend.

Specification gaming. Satisfying the stated objective in an unintended way.

Goal misgeneralisation. Competently pursuing the wrong goal when conditions change.

Alignment faking. Behaving as trained while observed in order to avoid being changed.

Scheming. Covertly pursuing a goal that conflicts with the operator's.

Sycophancy. Telling people what they want to hear.

Sandbagging. Deliberately underperforming on an evaluation.

Evaluation awareness. A model's ability to tell that it is being tested.

Interpretability. Reading what a model computes from its internal activity.

Superposition. Representing more concepts than there are neurons by overlapping them.

Sparse autoencoder. A tool that separates overlapping activity into more interpretable features.

AI control. Protocols that keep a system safe even if the model is working against its operators.

Trusted and untrusted models. A weaker model believed unable to scheme, and a stronger one whose intentions are unverified.

Scalable oversight. Methods for supervising work that the supervisor cannot fully check.

Prompt injection. Instructions hidden in content that an agent reads.

Jailbreak. An input that gets a model to ignore its safeguards.

Safety case. A structured argument, with evidence, that a system is safe enough for a given use.

Evaluation

Benchmark saturation. The point at which the best systems solve most items and the benchmark stops separating them.

Contamination. Test items appearing in training data.

Canary string. A unique marker placed in test data to detect leakage.

Elicitation. The effort spent drawing out a model's full capability during testing.

Confidence interval. A range that expresses the uncertainty of a measured value.

Network

Mint account. The on-chain record that defines a token on Solana.

Authority. An account with power to mint, freeze or upgrade.

Multisig. An arrangement in which several signers must approve an action.

Token extension. An optional feature of the Token-2022 program that changes how a token behaves.