Blog
Engineering August 24, 2026 6 min read OpenAI

Harness Engineering (3): Make the running system legible to the agent

An agent that can edit code still cannot verify the product. Autonomous engineering requires applications, browsers, logs, metrics, and acceptance criteria that agents can inspect directly.

J

Jonathan

Founder

A common completion condition for a coding agent is: the code changed, tests passed, and a Pull Request exists.

Users do not experience a diff. They experience running software. A button can pass a component test and still be covered by another element. An endpoint can return the right data while adding two seconds to a critical journey. A fix can remove one exception and break a neighboring workflow.

If only a human can open the browser, inspect the logs, and check the dashboard, the agent has completed implementation but not the engineering loop. As code generation accelerates, human QA becomes the next bottleneck.

OpenAI made each worktree runnable as an isolated application instance, connected the Chrome DevTools Protocol to Codex, and exposed DOM snapshots, screenshots, navigation, logs, metrics, and traces. The agent could reproduce a bug, apply a fix, and inspect evidence that the behavior changed.

Agent autonomy is bounded by how much verifiable reality the environment exposes.

Runnable is not the same as observable

A start command answers only whether the application launches. To judge correctness, an agent needs three layers of evidence.

The first is interface state: whether a page loaded, an element appeared, an interaction changed the DOM, and the visual result makes sense.

The second is system behavior: which requests fired, whether a background job ran, how external state changed, and which service produced an error.

The third is quality: startup time, endpoint latency, critical traces, error rate, and resource boundaries.

Terminal output alone forces the agent to infer outcomes indirectly. Structured UI state, logs, metrics, and traces let it reason from evidence.

Every task needs a disposable verification environment

Shared development environments create interference: tasks mutate the same state, logs mix together, failures are hard to reproduce, and destructive actions are poorly isolated.

A safer shape gives each change an isolated, temporary, reproducible runtime:

Task
  ↓
Worktree or sandbox
  ↓
Application instance + test data
  ↓
Browser session + logs + metrics + traces
  ↓
Destroy everything after verification

The environment does not need to mirror all of production. It does need isolation, reproducibility, a stable address, and safe cleanup.

Browser tools should return evidence, not just clicks

click, type, and navigate are only the beginning. Product verification also needs:

  • a DOM or accessibility tree for precise state assertions
  • screenshots for layout, occlusion, color, and responsive behavior
  • console and network evidence for frontend errors and failed requests
  • replayable navigation steps for before-and-after comparison

DOM data and screenshots complement each other. Screenshots reveal visual defects but are weak for exact element state. DOM snapshots support assertions but miss layout problems.

A strong UI task produces two sets of evidence: the steps and image that reproduce the failure, and the same journey after the fix. OpenAI describes Codex recording a video of the failure and a second video of the resolution as part of an end-to-end feature workflow.

Observability must become an agent interface

Many teams already have logs, metrics, and tracing, but only through dashboards designed for people. The agent may know the platform exists without having a stable way to query it or isolate results for its own task.

Agent-readable observability asks:

  • Are logs structured and filterable by task, service, and request?
  • Do metrics have stable names and a query interface?
  • Can a trace connect a user action to backend work?
  • Do tools return the relevant evidence instead of thousands of raw lines?
  • Is telemetry isolated per temporary environment?

OpenAI let Codex query logs with LogQL and metrics with PromQL. Requirements such as “startup must finish within 800ms” or “no span in these journeys may exceed two seconds” became conditions the agent could verify itself.

With this environment in place, individual Codex runs regularly worked on one task for more than six hours, often while the human team was asleep. Long runtime alone is not the achievement; the important part is that the agent retained access to enough observable evidence to keep making progress.

The specific query language is not the point. Quality requirements become executable only when they map to machine-readable signals.

Acceptance criteria must map to observable signals

“Improve login” is not verifiable. “After valid credentials, the user reaches the dashboard with no console error and the critical request stays below the agreed latency” maps to UI state, console output, network activity, and metrics.

A useful task definition includes:

Initial state: environment and data
Actions: the user or system journey
Expected result: UI and external-state changes
Quality boundary: latency, errors, security, or resource limits
Evidence: tests, queries, screenshots, or traces

Start with one high-value journey. Make the agent launch an isolated instance, prepare deterministic data, drive the browser, inspect runtime signals, and attach evidence to the result. Expand only after that path is reliable.

The console, network, and task-template recommendations in this section are a practical extension of OpenAI’s application-legibility pattern. The original case explicitly describes DOM snapshots, screenshots, navigation, logs, metrics, traces, and recorded videos; the broader evidence checklist shows how a team can apply the same principle to its own product.

Agent legibility is not a model-specific accommodation. Structured logs, reproducible environments, explicit acceptance criteria, and replayable journeys also improve testing, incident response, onboarding, and collaboration.

Adapted from OpenAI’s Harness engineering: leveraging Codex in an agent-first world.

Harness Engineering series

harness-engineering coding-agents observability agent-qa