Super IntelligenceDocsHome

Engineering

Inference

Serving a model well is the difference between a research result and something people can use. The cost of a fixed level of capability has fallen fast, and the workload has changed: agents send long, repetitive contexts and wait on tools.

State of the field

Techniques. Modern serving rests on five methods.

MethodEffect
Key-value cachingAvoids recomputing attention for tokens already processed
Prompt cachingReuses a stable prefix across requests
Continuous batchingFills the accelerator as requests arrive and finish
Paged memory for the cacheReduces fragmentation and raises throughput
Speculative decodingA small model drafts and the large model verifies

Engines. Open engines include vLLM for servers and llama.cpp for local use. Hugging Face's Text Generation Inference is in maintenance mode, and its repository recommends newer engines.

Cost. The price of GPT-3.5-level performance fell from $20 to $0.07 per million tokens between November 2022 and October 2024.

Open problems

  • Agent workloads fit poorly. Long cached contexts and bursts of waiting on tools are handled badly by schedulers built for batches.
  • Cost per token misleads. Reasoning tokens make the cost of a finished task diverge from the price per token.
  • Optimisations change behaviour. Quantisation and speculative decoding can shift outputs slightly, and equivalence is rarely tested.
  • Latency against throughput. The trade-off is sharp when many users share hardware.

Structure prompts for caching

Put the stable parts first: system instructions, then tool definitions, then the conversation. Track the cache hit rate as a primary metric.

Measure cost per successful task

cost per successful task = total spend on all attempts / number of tasks completed correctly

A cheaper model that fails more often can cost more per result.

Start from an open baseline

Before optimising, measure a standard engine with default settings.

vllm serve <model-id>

This starts a server that speaks the common chat-completions interface.

Evaluate what is served

Run the evaluation suite against the exact serving configuration: quantisation level, sampling parameters and speculative settings. A model evaluated at full precision and served at four bits has not been evaluated.