Context and Memory Management (C)

This page corresponds to §5 of Agent Harness Engineering: A Survey. Context is where prompt engineering and harness engineering meet. The engineering question is not “how do we fit more tokens?” It is “what should the model see at this step, and what should stay outside the window?”

Why Context Must Be Engineered

The paper gives three reasons large windows do not solve memory:

  • Quadratic attention cost: transformer attention relates token pairs; longer contexts remain structurally expensive even when kernels reduce constants.
  • Lost in the middle: information location matters. Relevant evidence in the middle of a long context is recovered worse than evidence near the beginning or end.
  • Context rot: frontier models degrade as input grows, often before the nominal window is full. Ambiguous semantic queries degrade more steeply than exact-match retrieval.

Context is a scarce resource even when the nominal window is large.

From Prompt Engineering To Context Engineering

Prompt engineering optimizes a mostly static input to one model call. Context engineering manages the full visible state of a multi-step run: system prompt, tool definitions, schemas, message history, tool results, retrieved documents, memory, scratch state, and dynamic work status.

All of those compete for the same attention budget.

Three Horizons

HorizonMechanismsWhat it solves
Short-term active windowPrompt altitude, token-efficient tool design, just-in-time retrieval, progressive disclosure, KV-cache-aware orderingKeeps the current step high-signal.
Mid-term session stateStructured notes, planning files, cross-run injectionPreserves progress across context resets without a full memory system.
Long-term persistent memoryMemGPT, Generative Agents, MemoryBank, Mem0, A-MEM, Hindsight, Honcho, cqSupports cross-session recall, indexing, forgetting, and shared user/task models.

The survey links these to academic memory taxonomies: short-term working memory and long-term semantic, procedural, and episodic memory; or, in another framing, a write-manage-read loop.

Long-Horizon Techniques

For 100+ turn agents, the paper highlights:

  • compaction: summarize accumulated state and restart the window; tune for recall before precision;
  • tool-result clearing: replace already-consumed verbose outputs with compact references;
  • subagent context isolation: delegate a subtask to a fresh window and return a 1000-2000 token synthesis;
  • hybrid decision rules: preload what is always needed, retrieve what is conditionally needed, compact near saturation, and fork subagents when exploration would pollute the main window.

The paper’s responsibility boundary is important: context management should be infrastructure responsibility, not a task the agent invents ad hoc on every run.

Context Drift

Context rot is a single-step phenomenon: more tokens in one inference can degrade reasoning. Context drift is a long-trajectory phenomenon: after many turns, the agent repeats completed work, contradicts itself, loses the original motivation, or preserves a subtly wrong summary.

Compaction and memory mitigate drift but do not solve it. Small summary errors accumulate. Current architectures often cannot detect that the internal task state has diverged from the true task state.

Open Problem: Context As State Estimation

The survey’s strongest reframing is to treat context management as state estimation. Future systems need:

  • uncertainty-aware summaries;
  • provenance for remembered facts;
  • contradiction handling;
  • explicit staleness markers;
  • recovery procedures that reconstruct missing state from durable artifacts rather than trusting compressed history.

Memory systems should be evaluated not only by recall accuracy, but by whether they prevent downstream action errors in multi-session tasks.

Was this page helpful?