Overview

ETCLOVG focuses on the harness around the model: execution, tooling, context, lifecycle, observability, verification, and governance. This section looks from the other side. It asks where model capabilities come from, and how those capabilities change the harness we need to build.

This is not just a theoretical distinction. Many harness components exist because the current model is not yet reliable at something important. If the model struggles to decompose a task, we add planning outside it. If tool use is brittle, we add routing, validation, and retries. If failure does not naturally lead to adjustment, we add reflection, search, and task logs. If preferences or norms are unstable, memory, rewards, or training enter the system.

As models improve, those responsibilities get redistributed. Some scaffolding becomes part of the model’s default behavior. Some judgments move in the opposite direction: they cannot be left to the model’s own discipline and must sit in verification, observability, and governance. The point of this section is not to ask whether model-side work can replace the harness. It is to ask: for the model we are deploying today, should this responsibility be handled by the model, the prompt, or the external system?

What This Section Answers

This is not a list of papers or a training manual. Each page is organized around an engineering question: how does this model-side method change the design of an agent system?

QuestionPageHarness Implication
Do explicit intermediate steps improve multi-step reasoning?Chain-of-ThoughtPrompts can carry some reasoning structure; production systems usually need summaries and result checks, not full reasoning traces.
How does reasoning connect to the outside world?ReActThe thought / action / observation loop becomes the basic structure of tool-using agents.
How can failure improve the next attempt?ReflexionFailed trajectories can become verbal feedback for memory, retries, task logs, and self-improvement loops.
Should the task be planned before it is solved?Plan-and-SolveExplicit planning helps decide whether the plan belongs in a prompt, a workflow, or a separate planner.
When one path is not enough, can the agent explore several?LATSAgent execution can be designed as evaluable, backtrackable search that depends on planning, evaluators, and state management.
Which behaviors can training turn into default capability?Training And Post-TrainingPretraining provides base capability; SFT, instruction tuning, tool training, preference optimization, and verifiable rewards shape default behavior after that.

Reading Path

Read the section in three steps.

First, structure a single run. Chain-of-Thought, Plan-and-Solve, and ReAct do not update model weights. They change how information is arranged inside one call, how the task is broken down, and how actions continue from observations. They ask how much capability can be unlocked with prompts and tool loops alone.

Second, keep exploring after failure. Reflexion and LATS do not treat one generation as the endpoint. They bring trajectories, failure, evaluation, and backtracking into the process. At this layer, the method is already close to harness design: the system must store attempts, diagnose failures, choose the next path, and decide when to stop.

Third, train capability into the model. Pretraining and post-training change the model itself, turning some behaviors that used to depend on prompts, tool examples, preference comparisons, or environment feedback into default capability. That does not make harnesses disappear. The more autonomously a model can act, the more the external system must provide permission boundaries, observability, verification, and rollback.

Boundary With Harnesses

Model-side methods answer what the model can learn. Harnesses answer how those capabilities are limited, combined, checked, and recovered in a real system.

That boundary moves as models change.

When the model is weaker, the harness often carries more scaffolding: explicit plans, strict workflows, external evaluators, more retries. As the model improves, some scaffolding can be deleted, but the risk can also grow because the model can take more consequential actions. The harness then shifts from “teach the model how to think” toward “limit what it can do, prove it did the right thing, and recover when it fails.”

Use this section to keep asking four questions:

  • Does the current model already have this capability reliably?
  • If not, should we compensate with a prompt, workflow, evaluator, memory, or training?
  • If a future model internalizes this capability, which harness scaffolds can be removed?
  • After removing scaffolding, do verification, audit, and safety boundaries still need to remain?

That is the relationship between model-side foundations and ETCLOVG: one explains where capability comes from; the other explains how production systems constrain that capability. The real design space for agent systems sits between the two.

Was this page helpful?