Blog
Engineering August 24, 2026 5 min read OpenAI

Harness Engineering (4): Turn engineering rules into executable constraints

Agents do not follow a rule forever because they read it once. A reliable harness turns architecture, quality requirements, and engineering taste into checks the system can enforce.

J

Jonathan

Founder

Documentation can explain how an agent should work. It cannot guarantee the agent will work that way every time.

Write “the service layer must not import the UI layer” in AGENTS.md, and the agent may follow it on most tasks. Under pressure to make one fix, it may still choose the shortest local path. Review catches the violation, but if the lesson remains a comment, the next agent can repeat it.

OpenAI took a stricter approach: do not micromanage every implementation step; encode the invariants that truly matter in custom lint and structural tests. The agent can choose the local implementation, but it cannot cross the architecture boundary.

Use prose to explain intent. Use code to enforce boundaries.

Repeated review comments reveal a missing constraint

Teams often treat review as the end of quality control. An engineer flags inconsistent naming, missing log fields, or unparsed input; the author fixes the change; the PR merges. The instance is fixed, but the judgment is not reusable.

At agent throughput, that approach stops scaling. The same comment appears in ten Pull Requests, and senior engineers repeat the same explanation.

Classify review feedback into three types:

  • Local judgment: a one-off product or implementation tradeoff that stays in human review.
  • A principle that needs explanation: useful across a class of work but not fully machine-decidable; document it.
  • An enforceable invariant: a violation is never acceptable; encode it in types, lint, tests, or CI.

If the same comment appears repeatedly, stop fixing only the current diff and ask whether the harness should own the rule.

Architecture becomes a boundary only when it can be checked

Architecture diagrams are easy to draw. Dependency direction drifts during ordinary work.

OpenAI divided each business domain into fixed layers and allowed dependencies to move in one direction. Cross-cutting concerns such as authentication, connectors, telemetry, and feature flags entered through an explicit Provider boundary. Structural tests rejected all other edges.

An existing project can start with a few questions:

  • Which directories may depend on each other?
  • Which external inputs must be parsed at the boundary?
  • Which modules can only be imported through a public entry point?
  • Which cross-cutting capabilities require a standard interface?
  • Which file-size, naming, logging, or reliability rules are non-negotiable?

Anything that can be decided statically should move into tooling. A boundary discoverable only by reading prose will leak as generation speed rises.

Error messages are part of the agent interface

A traditional lint failure may report only a rule name. A human can search documentation or ask a colleague. An agent can search too, but every search consumes time and context.

An agent-facing check should answer:

What rule was violated?
Why does it exist?
Which repair paths are allowed?

Instead of only service-must-not-import-ui, return an explanation: move shared types into the type layer, inject the capability through the Provider interface, and link the relevant architecture rule.

The checker becomes not only a gate, but also a teaching interface inside the current run.

Clear central constraints create local freedom

Encoding rules does not mean prescribing every implementation detail. Requiring one library, one loop shape, or one file pattern everywhere creates fossilized scaffolding.

Durable constraints target outcomes and boundaries:

  • Parse external data at the boundary without mandating a validation library.
  • Require structured logs and trace IDs without dictating function internals.
  • Require domains to communicate through public interfaces while allowing internal refactors.
  • Require verification evidence without forcing every task into the same test shape.

The center protects correctness, repeatability, and architecture. Inside those boundaries, the agent can choose the implementation.

Turn human taste into system capability

A practical progression is:

  1. Explain the issue in review.
  2. When it repeats, document the principle.
  3. Encode the decidable part in lint, tests, or a grader.
  4. Put repair guidance and documentation links in the error.
  5. Revisit the rule as models and the system evolve; remove obsolete scaffolding.

Harness engineering does not eliminate human taste. It lets valuable judgment apply repeatedly without requiring a person to restate it in every Pull Request.

Adapted from OpenAI’s Harness engineering: leveraging Codex in an agent-first world.

Harness Engineering series

harness-engineering coding-agents architecture guardrails