Engineering
Inference
Serving a model well is the difference between a research result and something people can use. The cost of a fixed level of capability has fallen fast, and the workload has changed: agents send long, repetitive contexts and wait on tools.
State of the field
Techniques. Modern serving rests on five methods.
| Method | Effect |
|---|---|
| Key-value caching | Avoids recomputing attention for tokens already processed |
| Prompt caching | Reuses a stable prefix across requests |
| Continuous batching | Fills the accelerator as requests arrive and finish |
| Paged memory for the cache | Reduces fragmentation and raises throughput |
| Speculative decoding | A small model drafts and the large model verifies |
Engines. Open engines include vLLM for servers and llama.cpp for local use. Hugging Face's Text Generation Inference is in maintenance mode, and its repository recommends newer engines.
Cost. The price of GPT-3.5-level performance fell from $20 to $0.07 per million tokens between November 2022 and October 2024.
Open problems
- Agent workloads fit poorly. Long cached contexts and bursts of waiting on tools are handled badly by schedulers built for batches.
- Cost per token misleads. Reasoning tokens make the cost of a finished task diverge from the price per token.
- Optimisations change behaviour. Quantisation and speculative decoding can shift outputs slightly, and equivalence is rarely tested.
- Latency against throughput. The trade-off is sharp when many users share hardware.
Recommended practice
Structure prompts for caching
Put the stable parts first: system instructions, then tool definitions, then the conversation. Track the cache hit rate as a primary metric.
Measure cost per successful task
cost per successful task = total spend on all attempts / number of tasks completed correctly
A cheaper model that fails more often can cost more per result.
Start from an open baseline
Before optimising, measure a standard engine with default settings.
vllm serve <model-id>
This starts a server that speaks the common chat-completions interface.
Evaluate what is served
Run the evaluation suite against the exact serving configuration: quantisation level, sampling parameters and speculative settings. A model evaluated at full precision and served at four bits has not been evaluated.