Engineering
Agents
An agent is a model running in a loop with tools. Most of the engineering difficulty is not in the loop. It is in what surrounds it: which tools exist, how context is managed over hours, and how the work is checked.
State of the field
Tools. The Model Context Protocol is the common interface between models and tools or data. Its 2026 specification makes the core stateless.
Long-running work. Agents that run for hours rely on managing context: compacting history, keeping structured notes in files, and using sub-agents that work in their own context and return a summary.
Reliability. The length of task completed at 50% reliability is rising quickly, and reliability at higher thresholds lags far behind. See Reasoning and autonomy.
The loop
The whole control flow fits in a few lines. This example uses the Anthropic Python SDK. TOOLS and run_tool are supplied by the application, and run_tool must execute inside a sandbox.
import anthropic
client = anthropic.Anthropic()
messages = [{"role": "user", "content": task}]
while True:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
tools=TOOLS,
messages=messages,
)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
break
results = [
{
"type": "tool_result",
"tool_use_id": block.id,
"content": run_tool(block.name, block.input), # sandboxed
}
for block in response.content
if block.type == "tool_use"
]
messages.append({"role": "user", "content": results})
Open problems
- Reliability over many hours. Errors compound, and state drifts after history is compacted.
- Model or harness? The same model scores very differently under different agent software. Leaderboards mix the two.
- Coordination cost. Several agents use more tokens, duplicate work and make conflicting edits.
- Memory. It needs to be durable, auditable and resistant to poisoning.
Recommended practice
Start simple
Begin with one loop and a few well-documented tools. Add sub-agents only for subtasks that can run in parallel and need a lot of context.
Keep state outside the context window
| State | Where it lives |
|---|---|
| Progress | A notes file the agent updates |
| Work product | Commits in version control |
| Correctness | Test results |
| Decisions | The audit log |
Any fresh context should be able to resume from these alone.
Prefer verification to self-assessment
Give the agent tools that check its work: tests, type checkers, linters, screenshots. Do not accept an agent's own statement that a task is done.
Report the harness with the score
Every agent result states the harness, the tools, the token budget and the number of attempts.
Design for interruption
A person must be able to pause, inspect and redirect a running agent. See Safety engineering.