Why AI agents need a harness, not just a better model
Better models raise the ceiling, but reliable agents come from the system around the model: context, tools, constraints, verification, correction, and observable loops.
aibuddy Team
Runtime
The model is not the whole agent
It is tempting to explain every agent failure as a model problem. The model chose the wrong tool, forgot the goal, stopped too early, wrote malformed JSON, or confidently claimed an action had succeeded when it had not. A stronger model will fix some of those failures. It will reason better, follow instructions more closely, and need fewer hints.
But that is not the same as saying the model is the agent.
An agent is a model running inside a system that gives it context, tools, state, boundaries, feedback, and a way to keep going. Anthropic draws a useful distinction here: workflows use LLMs and tools along predefined code paths, while agents let the LLM dynamically direct its own process and tool use. OpenAI’s Agents SDK makes a similar engineering move: you can call the Responses API directly when you want to own the loop, tool dispatch, and state yourself, or use the SDK when you want a runtime to manage turns, tools, guardrails, handoffs, sessions, and traces.
That runtime is what I mean by a harness. It is not a magic framework, and it is not a pile of prompts. It is the engineering layer that turns model calls into a product that can act, observe, recover, and be trusted.
A harness is context, tools, and reliability around them
The simplest formula is still useful:
Agent = LLM + Context + Tools
But a production agent needs the expanded version:
Agent = Model + Harness
Where the harness is:
Context + Tools + Constrain + Verify + Correct
Context lets the model perceive: instructions, conversation history, documents, retrieved knowledge, tool definitions, memory, and the current state of the task. Tools let the model act: call APIs, search, edit files, run code, create tickets, or hand work to another agent.
Those two pieces make the agent capable. They do not make it reliable.
Reliability comes from the outer layers. Constrain decides what the agent may do: permissions, schemas, sandbox boundaries, rate limits, spending limits, human approval gates. Verify checks whether the action actually worked: database state, tests, assertions, graders, traces, tool outputs. Correct decides what happens when something goes wrong: retry, ask for clarification, roll back, compact context, resume from durable state, or stop safely.
This is the real gap between a demo and a product. A demo shows that the model can do something once. A harness makes it possible to do the right thing repeatedly, under messy conditions, with enough evidence to debug when it fails.
The loop matters more than the label
Not every useful system should be a free-running autonomous agent. Anthropic’s guidance is blunt and correct: start with the simplest solution, add agentic complexity only when it improves outcomes, and prefer predictable workflows when the path is known.
That means harness engineering is not “always build a giant agent framework.” Often the right answer is a simple workflow: classify, route, call a tool, check the result. Sometimes it is prompt chaining with a gate after each step. Sometimes it is an evaluator-optimizer loop. Sometimes, when the number of steps cannot be known in advance, it really is an agent that plans, calls tools, observes the environment, and adapts.
The important question is not whether you call it a workflow, an agent, or an orchestration layer. The important question is who owns the loop.
Something has to decide what the model sees next, which tool result enters the transcript, when to ask a human, when to stop, and how to preserve state across turns. If you use a managed SDK, it owns some of that. If you build directly on the model API, you own more of it yourself. Either way, those choices are part of the product, not an implementation detail.
Tool use is an interface design problem
The model does not execute tools. It requests tool use. The harness parses the request, validates it, applies permissions, executes it in the right environment, and returns the result.
That boundary is where a lot of agent quality is won or lost.
Anthropic calls this ACI: Agent-Computer Interface. A good tool is not just a thin wrapper around an internal API. It is an interface designed for a model to understand and use correctly. Names should be obvious. Parameters should be hard to misuse. Tool descriptions should include constraints, examples, and edge cases. If relative paths cause errors after the agent changes directories, require absolute paths. If a refund must never exceed the order total, make that impossible at the tool boundary instead of hoping the model remembers.
OpenAI’s Agents SDK reflects the same direction: function tools generate schemas, MCP tools are exposed through a consistent interface, tool guardrails can run checks before and after custom tool execution, and traces record tool calls alongside model turns and handoffs.
The lesson is simple: don’t just prompt the model harder. Design the environment so the correct action is easier than the wrong one.
Long-running work needs durable state
Short tasks can often fit in one model call or one context window. Long-running agents cannot rely on that.
Anthropic’s long-running harness work makes this concrete. Agents that work across hours or days have to span multiple context windows. Compaction helps, but it is not enough by itself. A later session still needs to know what happened, what was attempted, what remains, and which artifacts are authoritative. Otherwise the next model instance is like a new engineer joining mid-shift with no handoff notes.
For product systems, this means the harness needs durable state outside the model call. Keep an append-only message record or trace. Preserve tool results. Store artifacts. Summarize older context without erasing the anchors the agent needs. Make resumability a property of the system, not a lucky side effect of a long context window.
This is also why observability is not optional. OpenAI traces record model generations, function calls, guardrails, handoffs, and custom events. Anthropic’s eval terminology similarly treats transcripts, outcomes, graders, and evaluation harnesses as first-class objects. If you cannot see the trajectory, you cannot evaluate the agent. If you cannot evaluate it, you cannot safely improve it.
Scaffolding should evolve as models improve
There is a danger in harness work: scaffolding can fossilize.
Every workaround encodes an assumption about what the current model cannot do. It cannot plan well, so you pre-split every task. It cannot manage context, so you reset too aggressively. It cannot use a certain tool format, so you wrap it in extra ceremony. Some of those assumptions are true today. Some will become false when the model improves.
Anthropic’s Managed Agents work is a useful correction to the old “more scaffolding is always better” instinct. Harnesses encode assumptions, and those assumptions need to be re-questioned. The durable part is not every workaround. The durable part is the interface: session, harness, sandbox, tools, permissions, traces, evals.
So build harnesses out of primitives, not habits. Keep the message record stable. Keep tool boundaries explicit. Keep the sandbox replaceable. Keep checks measurable. Delete scaffolding when it starts fighting the model instead of helping it.
Better models raise the ceiling; harnesses raise the floor
The right answer is not “models don’t matter.” They matter enormously. Model choice affects reasoning, tool use, latency, cost, context length, and multimodal capability. For many agent tasks, a stronger reasoning model is the difference between viable and broken.
But model progress and harness engineering solve different parts of the problem.
A better model raises the ceiling of each decision. A better harness raises the floor of the whole run. It makes context available, tools usable, permissions enforceable, outcomes verifiable, failures recoverable, and behavior observable.
That is why agents need a harness. Not because models are weak, and not because every team should adopt a heavyweight framework. Agents need a harness because real work happens across turns, tools, state, failures, and trust boundaries. The model supplies intelligence. The harness turns that intelligence into reliable action.
Related reading
- Harness Engineering — the fuller docs version of context, tools, constrain, verify, and correct.
- Context engineering is the bottleneck in long agent runs — why the context window, not the model alone, sets the ceiling on long agent runs.