Context engineering is the bottleneck in long agent runs
Long agent runs fail when the model sees the wrong slice of state. The hard part is not stuffing the window; it is managing context as working memory across turns, tools, compaction, memory, and cache.
aibuddy Team
Runtime
The window is working memory, not storage
Every model call is a fresh inference over the input you send. The model can use databases, files, search results, memories, and tool outputs only after the harness has retrieved them and placed the relevant pieces into context. That makes the context window less like a hard drive and more like working memory: bounded, temporary, and easy to pollute.
This is the first trap in long agent runs. Teams hear “larger context window” and assume they can keep everything. In practice, more context is not automatically better. Low-signal tokens compete with high-signal tokens. Old tool output, repeated logs, stale hypotheses, and irrelevant documents do not merely cost money; they dilute the model’s attention and make the useful facts harder to find.
Anthropic frames context engineering as choosing the configuration of context most likely to produce the desired behavior. Our docs put the same idea more bluntly: model capability is the foundation, context sets the ceiling. A strong model with poor context behaves like a brilliant new hire dropped into a company with no documentation, no process, and decisions hidden in private chats.
So the core question is not “how much can we fit?” It is:
What is the smallest high-signal view of the task the model needs on this turn?
The bottleneck is the context lifecycle
A real agent is not one prompt. It is a loop.
The model reads context, proposes an action, the harness executes tools, the results come back, and the next model call starts from a new context view. That view includes a stable head, tool definitions, message history, current task state, retrieved memories, and whatever the last step produced. It grows with every turn.
If the runtime only appends, the run eventually hits two limits. The obvious one is the hard window limit. The subtler one arrives earlier: context rot. Even before the window is full, the model’s ability to recall and use specific information can degrade as the window gets crowded.
This is why long-running agents need context management, not just better prompts. The runtime has to decide, at every model-call boundary:
- what stays raw
- what gets summarized
- what moves to memory or files
- what can be reloaded on demand
- what should be isolated in a subagent
- what must stay byte-stable for prompt caching
Those are product decisions and architecture decisions. They decide whether the agent keeps coherence over a two-hour run or forgets the original goal at step forty.
Compaction is not truncation
The easiest way to shrink context is to delete old messages. It is also the wrong default.
The beginning of a task often contains the goal, constraints, user preferences, and early decisions. The most recent turns often contain raw scrollback: long tool results, partial logs, or exploratory output that has already served its purpose. A simple “drop oldest first” policy can delete the load-bearing context and keep the noise.
Compaction is different. It turns raw history into a smaller derived view. The durable record should remain append-only; what changes is the projection sent to the model. That distinction matters. The user and system keep the full history, retries can re-derive the same view, and anything compacted out of the model’s immediate window can still be recovered from the source record or artifacts.
Good compaction is graded. First shrink large tool results into the lines or identifiers that mattered. Then summarize completed arcs of work. Only under real pressure should the runtime fold large spans of history into a dense task summary. And even then, it should protect the recent window, because the last few steps are usually what the model needs to continue coherently.
The summary should preserve conclusions, not transcript flavor: the user’s request, standing constraints, decisions and rationale, deliverable paths, verification status, unresolved TODOs, and approaches already ruled out. Process can usually go. Anchors cannot.
Memory and notes are what must survive the window
Some facts should not depend on the rolling transcript.
The user’s goal. A rule stated once at the start. A project convention discovered through a file search. A decision made after comparing two approaches. A TODO that must not silently revert. If these live only in the raw dialogue, compaction will eventually put them at risk.
This is where memory and structured notes fit. They are not the same as transcript. Transcript is chronological evidence. Memory is durable knowledge. Structured notes are working memory externalized to files or state: current plan, decisions, open questions, verified facts, and progress so far.
Anthropic’s long-running agent guidance makes the same point with structured notes: agents that write and read state files can keep task continuity across context resets. Our docs describe a related “status bar” pattern: the harness turns scattered implicit state into a small explicit block near the end of context, such as current TODOs, attempt counts, elapsed time, or environment status.
The pattern is simple: do not make the model rediscover state from raw scrollback every time. Compute or record the state once, make it inspectable, and inject the relevant part when it matters.
Cache is the cost side of the same design
Context engineering is also economics.
A long agent run repeatedly sends a large prefix: system instructions, tool definitions, project context, message history, and retrieved state. Providers can cache repeated prefixes, but only under specific rules. OpenAI applies prompt caching automatically for sufficiently long repeated prefixes and reports cached tokens in usage. Anthropic supports automatic caching and explicit cache breakpoints; exact matching and cache lifetime matter, and explicit placement can be important for long conversations.
The design implication is the same across providers: stable content belongs early; volatile content belongs late. Do not put timestamps, session-specific state, random ordering, or constantly changing reminders into the system head. Append dynamic state near the tail. If compaction rewrites a cached prefix too frequently, you save window space but lose cache efficiency. If you never compact, the prefix stays cacheable but grows until it becomes too large or too noisy.
So cost is not controlled by one magic number. It is the product of two levers:
- token count: managed by pruning, isolation, memory, and compaction
- price per input token: improved by prompt caching and byte-stable prefixes
Cache hit rate is the headline metric on the price lever. Context occupancy and signal density are the headline metrics on the size and quality lever. You need both.
The real skill is choosing what the model sees now
Compaction, memory, status blocks, subagents, tool-result clearing, and prompt caching are not separate tricks. They are all answers to the same question:
What should be in the model’s working memory for this next decision?
Sometimes the answer is raw recent history. Sometimes it is a summary. Sometimes it is a file path, not the file. Sometimes it is a structured note. Sometimes it is nothing, because the work should happen inside a subagent and return only a conclusion.
That is why context engineering is the bottleneck in long agent runs. Better models matter; they raise the ceiling of what can be done with a good context view. But the view itself is still engineered. If you give the model the wrong slice of state, it will reason beautifully over the wrong problem. If you give it the right slice, even a smaller model can often do the work.
The agent does not need everything. It needs the right thing, at the right time, in a form it can use.
Sources
- Effective context engineering for AI agents — Anthropic’s framing of context as a finite, diminishing-return resource.
- Context engineering: memory, compaction, and tool clearing — Anthropic cookbook on how memory, compaction, and clearing compose.
- Prompt caching — Claude API docs on automatic caching, explicit breakpoints, exact-prefix matching, and TTLs.
- Prompt Caching in the API — OpenAI’s automatic prompt caching behavior and cached-token reporting.
Related reading
- Why AI agents need a harness, not just a better model — the coordination problems a better model does not solve for you.