Blog
Engineering August 24, 2026 7 min read OpenAI

Harness Engineering (2): Give the agent a map of the repository

A giant AGENTS.md does not solve context. Agents need a short entry point, layered knowledge, and a system of record that stays aligned with the code.

J

Jonathan

Founder

An agent entering an unfamiliar repository resembles a new engineer. It can read code and search files, but it does not know which directories carry the business, which constraints came from past incidents, which documents are stale, or why an odd-looking implementation was an intentional tradeoff.

The first response is often a larger AGENTS.md: project overview, repository layout, coding rules, commands, product principles, and every warning in one file.

OpenAI tried that approach and abandoned it. The team eventually treated a roughly 100-line AGENTS.md as a table of contents, not an encyclopedia. Detailed knowledge lived in versioned architecture documents, product specifications, execution plans, and a structured docs/ tree.

Give the agent a map it can follow, not a manual it must memorize on every run.

A giant AGENTS.md wastes both context and attention

Putting all guidance in one file creates four problems.

First, context is finite. A large instruction file competes with the task, relevant code, and tool results. The agent sees more text but may miss what matters now.

Second, priority disappears. Global invariants, naming preferences, deployment warnings, and local component conventions become indistinguishable. If everything is marked critical, nothing is.

Third, the file decays. New rules are added; old rules are rarely removed. Stale guidance is often more dangerous than missing guidance because it looks authoritative.

Fourth, one large blob is difficult to validate. It is hard to check ownership, freshness, coverage, links, and whether the described behavior still exists.

The real question is not how long AGENTS.md should be. It is whether repository knowledge has clear boundaries, a stable entry point, and a maintenance loop.

The repository should become the engineering system of record

For an agent, inaccessible knowledge effectively does not exist.

A Slack discussion may explain why the team rejected a dependency. An incident may establish a retry constraint. A product meeting may explain an intentionally unusual interaction. If those decisions remain in chat or in people’s heads, the next agent cannot use them.

Not all company knowledge belongs in Git. The repository should contain the engineering facts that affect implementation and verification:

  • current architecture and allowed dependency directions
  • product-domain boundaries and key behavior
  • decisions and their rationale
  • active plans, completed work, and known debt
  • reliability, security, and quality requirements
  • executable development, test, and release procedures

Versioning these artifacts lets knowledge change with code, allows review to catch mismatches, and gives the agent one searchable environment for both implementation and intent.

Progressive disclosure keeps the working context focused

A practical structure can begin with three layers:

AGENTS.md       entry point, commands, global invariants, and navigation
ARCHITECTURE.md top-level domains, packages, boundaries, and deeper links
docs/
├── design-docs/    decisions and engineering principles
├── product-specs/  behavior and acceptance criteria
├── exec-plans/     active plans, completed plans, decision logs
├── generated/      schema and other generated facts
└── references/     project-specific external references

The names matter less than the loading strategy. The agent starts from a short, stable entry point, identifies the relevant domain, and follows links to deeper material only when the task requires it. Unrelated detail never enters the working context.

That is progressive disclosure: navigation first, expansion on demand.

Legibility changes technology choices

Once repository knowledge becomes the system of record, dependency choices are no longer judged only by developer convenience. OpenAI favored technologies and abstractions that an agent could inspect and reason about inside the repository. “Boring” technology often performed well because its APIs were stable, its behavior was composable, and it was well represented in model training data.

In one case, the team implemented a focused map-with-concurrency helper instead of adopting a generic p-limit-style package. The local implementation integrated directly with OpenTelemetry, had 100% test coverage, and matched the runtime’s expected behavior. The point is not to reimplement every dependency. It is to notice when opaque upstream behavior costs more agent effort than a small, fully inspectable local abstraction.

Plans and decisions are first-class engineering artifacts

Complex agent tasks span hours and sometimes multiple context windows. If the plan exists only in the first prompt or raw transcript, compaction and restarts can erase critical decisions.

A durable execution plan should record the goal and non-goals, current progress, standing constraints, key decisions and rationale, deferred debt, and the evidence required for completion. Small changes can use lightweight plans; larger changes need resumable artifacts.

Completed plans should remain discoverable. They explain why the code evolved into its current shape and let a later agent continue without reconstructing history from chat.

Documentation needs its own feedback loop

Moving knowledge into the repository is not enough. Unmaintained documentation simply creates a larger source of stale context.

Maintenance can operate at three levels:

  1. Structure checks: validate required files, links, indexes, owners, and metadata.
  2. Change checks: require documentation updates when code changes public behavior or architecture boundaries.
  3. Content gardening: regularly compare documents with the code and open small fixes for stale paths, invalid commands, mismatched schemas, or missing indexes.

The third layer is especially suitable for agents because the work is repetitive, bounded, and evidence-driven.

A repository map should answer five questions

An effective entry point lets the agent answer:

  1. What does this project do, and what is outside its scope?
  2. What are the core domains and dependency directions?
  3. Which commands develop, test, and verify the product?
  4. Which rules are global invariants, and what enforces them?
  5. Where should the agent look next for product, architecture, reliability, or security questions?

If an answer takes hundreds of lines, move the detail to a dedicated document and link it. If a rule is stable and important enough, promote it from prose into an automated check.

The repository map does not make the agent know everything. It lets the agent recognize what it does not know and find a trustworthy answer.

Adapted from OpenAI’s Harness engineering: leveraging Codex in an agent-first world.

Harness Engineering series

harness-engineering coding-agents context-engineering agents-md