Multi-Agent Architecture
Anthropic’s article tackles two connected problems: getting Claude to produce higher-quality frontend designs, and getting it to build complete applications without human intervention.
Earlier frontend design skills and long-running coding harnesses improved Claude well above baseline through prompt engineering and harness design, but both eventually hit ceilings. To push further, Anthropic drew inspiration from GANs: separate the agent that generates work from the agent that evaluates it.
The resulting architecture uses three roles: planner, generator, and evaluator. Together, they can produce richer full-stack applications across multi-hour autonomous coding sessions.
Why Naive Implementations Fall Short
Earlier long-running harness work used an initializer agent to decompose a product spec into a task list, then a coding agent to implement one feature at a time while handing off context through structured artifacts. This solved multi-session continuity, but complex tasks still drifted over time.
Anthropic observed two persistent problems.
First, models lose coherence on lengthy tasks as the context window fills. Some models also show context anxiety: they begin wrapping up prematurely as they approach what they believe is the context limit. Context resets address this by clearing the window, starting a fresh agent, and using a structured handoff to carry state and next steps.
This differs from compaction. Compaction summarizes earlier conversation in place so the same agent can continue, but it does not give the agent a clean slate, so context anxiety can persist. Reset gives the model a clean slate at the cost of requiring a strong handoff artifact.
Second, agents are poor self-evaluators. When asked to evaluate their own work, they tend to confidently praise mediocre output. This is especially pronounced in subjective tasks like design, where there is no binary software-test equivalent. Even on verifiable tasks, poor judgment can impede completion.
Separating the agent doing the work from the agent judging it is a strong lever. The separation does not immediately eliminate leniency because the evaluator is still an LLM. But tuning a standalone evaluator to be skeptical is much more tractable than making a generator critical of its own work.
Frontend Design: Making Subjective Quality Gradable
Anthropic started with frontend design, where self-evaluation problems were most visible. Without intervention, Claude tends toward safe, predictable layouts that are functional but visually unremarkable.
Two insights shaped the design harness:
- Aesthetics cannot be fully reduced to a score, and taste varies, but grading criteria can encode design principles and preferences.
- Separating frontend generation from frontend grading creates a feedback loop that drives stronger outputs.
The generator and evaluator both received four criteria:
| Criterion | Focus |
|---|---|
| Design quality | Does the design cohere as a whole rather than a collection of parts? Do color, typography, layout, imagery, and details create a distinct mood and identity? |
| Originality | Is there evidence of custom decisions rather than template layouts, library defaults, or common AI-generated patterns? |
| Craft | Technical execution: typography hierarchy, spacing consistency, color harmony, contrast ratios. |
| Functionality | Usability independent of aesthetics: can users understand the interface, find primary actions, and complete tasks? |
Design quality and originality were emphasized over craft and functionality. Claude already performed well on craft and functionality by default, while design and originality were often bland. The criteria explicitly penalized generic “AI slop” patterns such as purple gradients over white cards and pushed the model toward more aesthetic risk-taking.
The evaluator also needed calibration. Anthropic used few-shot examples with detailed score breakdowns to align the evaluator’s judgment with human preferences and reduce score drift across iterations.
The loop ran on the Claude Agent SDK. A generator created an HTML/CSS/JS frontend from a prompt. The evaluator, equipped with Playwright MCP, interacted with the live page, took screenshots, studied the implementation, scored each criterion, and wrote a critique. That feedback fed the next generation.
Runs usually took 5 to 15 iterations. Because the evaluator interacted with the page rather than just grading a static screenshot, full runs could take hours. After each evaluation, the generator decided whether to refine the current direction or pivot to a different aesthetic.
Key observations:
- Evaluator assessments improved over iterations before plateauing.
- Criteria wording shaped outputs in unexpected ways; phrases like “museum quality” pushed designs toward visual convergence.
- Improvement was not perfectly linear; middle iterations were sometimes preferable to final ones.
- Implementation complexity increased as the generator responded to feedback.
- Even first iterations beat the no-prompting baseline, showing that the criteria language itself steered the model away from generic defaults.
One notable example was a Dutch art museum website. By iteration nine, Claude had produced a clean dark-themed landing page. On iteration ten, it scrapped the approach and reimagined the site as a spatial experience: a CSS-perspective 3D room, checkered floor, artworks placed freely on walls, and doorway-based navigation between gallery rooms.
Scaling To Full-Stack Coding
With the frontend results in hand, Anthropic applied the GAN-inspired pattern to full-stack development. The generator-evaluator loop maps naturally to the software development lifecycle, where code review and QA play the evaluator role.
Architecture
The earlier long-running harness solved multi-session coding with an initializer agent, a one-feature-at-a-time coding agent, and context resets. Context resets were important because Sonnet 4.5 showed context anxiety.
With Opus 4.5, that behavior was largely gone, so this harness dropped context resets. Agents ran as one continuous session across the build, with the Claude Agent SDK’s automatic compaction handling context growth.
The system used three agent personas.
Planner
The earlier harness required the user to provide a detailed spec upfront. The planner automates that step by expanding a simple 1-4 sentence prompt into a full product spec.
The planner is prompted to:
- Be ambitious about scope.
- Focus on product context and high-level technical design.
- Avoid granular technical implementation details.
- Find opportunities to weave AI features into the product.
Avoiding granular implementation is intentional. If the planner specifies a wrong low-level detail early, that mistake cascades downstream. It is better to constrain deliverables and let later agents determine the path.
Generator
The generator carries forward the one-feature-at-a-time approach. In the first version, it worked in sprints, picking one feature from the spec at a time.
The implementation stack was React, Vite, FastAPI, and SQLite, later PostgreSQL. At the end of each sprint, the generator self-evaluated before handing work to QA and used git for version control.
Evaluator
Earlier harness outputs often looked impressive but had real bugs when used. The evaluator catches those. With Playwright MCP, it clicks through the running app like a real user, testing UI features, API endpoints, and database state.
It then grades each sprint against criteria adapted from the frontend experiment: product depth, functionality, visual design, and code quality. Each criterion has a hard threshold. If any criterion falls below threshold, the sprint fails and the generator receives detailed feedback.
Sprint Contract
Before each sprint, the generator and evaluator negotiate a sprint contract: what “done” means for that chunk of work before code is written.
This bridges the gap between high-level product stories and testable implementation. The generator proposes what it will build and how success will be verified. The evaluator reviews whether the right thing is being built. They communicate through files until they agree.
This keeps work faithful to the spec without over-specifying implementation too early.
Running The Harness: Retro Game Maker
The first harness version used Claude Opus 4.5. Anthropic compared the full harness with a single-agent system on the same prompt:
Create a 2D retro game maker with features including a level editor, sprite editor, entity behaviors, and a playable test mode.
| Harness | Duration | Cost |
|---|---|---|
| Solo | 20 min | $9 |
| Full harness | 6 hr | $200 |
The full harness was more than 20x more expensive, but the output quality difference was obvious.
The solo run looked plausible at first: an interface for constructing levels and components, then pressing play. But use revealed problems. The layout wasted space. Fixed-height panels left most of the viewport empty. The workflow was rigid and did not guide users to create sprites and entities before placing them. Most importantly, the game was broken: entities appeared on screen but did not respond to input. Code inspection showed the wiring between entity definitions and the game runtime was broken.
The full harness started from the same one-sentence prompt, but the planner expanded it into a 16-feature spec across ten sprints. The spec included not only the core editors and play mode, but also sprite animation, behavior templates, sound effects and music, AI-assisted sprite generation, AI level design, and export with shareable links.
The planner also read Anthropic’s frontend design skill and used it to define a visual design language. For each sprint, the generator and evaluator negotiated a contract with implementation details and testable behaviors.
The full harness output was more polished. The canvas used the full viewport, panels were sensibly sized, and the UI had a consistent visual identity. Some product-flow clunkiness remained; for example, users still had to discover that sprites and entities should be built before populating a level. That suggested a gap in the base model’s product intuition rather than something the harness had already solved.
The largest difference was play mode. The author could move an entity and play the game. Physics still had rough edges, and AI-generated levels could produce walls that trapped the player, but the core gameplay worked.
Logs showed that the evaluator kept implementation aligned with the spec. Each sprint, it walked through sprint contract criteria in Playwright and filed bugs for deviations. Sprint 3 alone had 27 criteria covering the level editor, and the evaluator’s findings were specific enough to act on without extra investigation.
Examples:
| Contract criterion | Evaluator finding |
|---|---|
| Rectangle fill tool allows click-drag to fill a rectangular area with selected tile | FAIL: the tool only places tiles at drag start and end, and fillRectangle is not triggered properly on mouseUp. |
| User can select and delete placed entity spawn points | FAIL: the Delete key handler requires both selection and selectedEntityId, but clicking an entity only sets selectedEntityId. |
| User can reorder animation frames via API | FAIL: PUT /frames/reorder is defined after /{frame_id}, so FastAPI parses reorder as an integer frame id and returns 422. |
QA Agent Challenges
Getting the evaluator to this level required work. Out of the box, Claude is a poor QA agent.
In early runs, it identified real issues, then talked itself into approving the work anyway. It also tested superficially instead of probing edge cases, so subtle bugs slipped through.
The tuning loop was:
- Read evaluator logs.
- Find examples where evaluator judgment diverged from human judgment.
- Update the QA prompt to address those failures.
After several rounds, the evaluator graded in a more reasonable way. Even then, the harness output showed limits: small layout issues, unintuitive interactions, and bugs in deeply nested features the evaluator did not exercise thoroughly.
Compared with the solo run, however, the lift was clear: the central feature worked.
Iterating On The Harness
The first results were encouraging, but the harness was bulky, slow, and expensive. The next step was to simplify without degrading performance.
The general principle: every harness component encodes an assumption about what the model cannot do on its own. Those assumptions should be stress-tested because they may be wrong, and they can go stale quickly as models improve.
The first simplification attempt cut too much and added new ideas, but failed to reproduce the original performance. It also became hard to tell which parts were load-bearing. Anthropic then moved to a more methodical approach: remove one component at a time and inspect the impact.
Opus 4.6 provided more motivation to simplify. It improved long-horizon planning, longer agentic tasks, reliability in larger codebases, code review, debugging, and long-context retrieval: all capabilities the harness had been supplementing.
Removing The Sprint Construct
The first removed component was the sprint construct. Sprints had decomposed work into chunks the model could handle coherently. With Opus 4.6, there was reason to believe the model could handle the job without this decomposition.
The planner and evaluator remained because they still added clear value:
- Without the planner, the generator under-scoped and started building from the raw prompt without a rich spec.
- The evaluator moved from per-sprint grading to a single pass at the end of the run.
As model capability increased, whether the evaluator was load-bearing became task-dependent. On 4.5, many builds sat close to the boundary of what the generator could do alone, so the evaluator caught meaningful issues. On 4.6, that boundary moved outward. Some tasks no longer needed the evaluator; others still benefited where the build remained at the generator’s capability edge.
The practical implication: the evaluator is worth its cost when the task is beyond what the current model can reliably do solo.
Anthropic also added prompting so the harness could build AI features into each app, especially by getting the generator to build a proper agent that could drive the app’s own functionality through tools. This took iteration because the relevant knowledge was recent and thinly represented in training data.
Updated Harness: Browser DAW
To test the updated harness, Anthropic used a Digital Audio Workstation prompt:
Build a fully featured DAW in the browser using the Web Audio API.
The run still took about four hours and around $124 in token costs. Most time went to the builder, which ran coherently for over two hours without the sprint decomposition Opus 4.5 had needed.
| Agent / Phase | Duration | Cost |
|---|---|---|
| Planner | 4.7 min | $0.46 |
| Build Round 1 | 2 hr 7 min | $71.08 |
| QA Round 1 | 8.8 min | $3.24 |
| Build Round 2 | 1 hr 2 min | $36.89 |
| QA Round 2 | 6.8 min | $3.09 |
| Build Round 3 | 10.9 min | $5.88 |
| QA Round 3 | 9.6 min | $4.06 |
| Total V2 Harness | 3 hr 50 min | $124.70 |
As before, the planner expanded the one-line prompt into a full spec. Logs showed the generator did well planning the app, designing the app’s agent, wiring it up, and testing before QA.
QA still caught real gaps. In the first round, it noted that the app had strong design fidelity, a solid AI agent, and good backend, but failed on feature completeness: several core DAW features were display-only, such as clips not being draggable on the timeline, missing instrument panels, and missing visual effect editors.
In the second round, QA caught more gaps: audio recording was still stub-only, clip edge-resize and split were missing, and effect visualizations were numeric sliders rather than graphical EQ curves.
The final app was not a professional music production tool, and Claude cannot hear audio, which limited QA for musical taste. But the app had the core pieces of a browser DAW: arrangement view, mixer, and transport. The integrated agent could set tempo and key, create melody and drums, adjust mixer levels, and add reverb through tools.
What Comes Next
As models improve, they will work longer and handle more complex tasks. Sometimes scaffolding around the model will matter less over time, and developers can wait for the next model to solve certain problems.
At the same time, stronger models create more room for ambitious harnesses that achieve tasks beyond baseline capability.
Lessons to carry forward:
- Experiment with the target model.
- Read traces on realistic problems.
- Tune performance toward desired outcomes.
- For complex tasks, decomposing work and applying specialized agents may still create headroom.
- When a new model lands, re-examine the harness, remove components that are no longer load-bearing, and add new ones that unlock previously impossible capabilities.
Anthropic’s conclusion is that the space of interesting harness combinations does not shrink as models improve. It moves. The AI engineer’s work is to keep finding the next useful combination.
Source
Harness Design for Long-Running Application Development — Prithvi Rajasekaran, Anthropic, March 2026.