Agent Team (Implementation)
Separate configuration, assignments, and runs
The design starts with a capable LLM that can choose useful delegation, communicate discoveries, and judge results. The harness makes those choices executable and observable. See Coordination model for the product principles and model-compatibility boundaries.
Four concepts have different lifetimes:
| Concept | Meaning |
|---|---|
| AgentTeam catalogue | Reusable team configuration and expert roles. |
| Channel roster | Agents available in a conversation; a roster does not itself start execution. |
| Runtime Team | A work session created by team_create, closed by team_dissolve. |
| Member / Task / Run | A persistent member owns one logical task and can execute multiple physical runs. |
A delivered member can be called back while its runtime team remains active. Persistence refers to identity and records; a wake run does not automatically restore the full previous conversation.
The chat host and members share the agent loop
runChatTurn builds the Lead’s runtime, tools, and instructions. When the system Lead has the team tool and other channel members, it receives TEAM_LEAD_OVERLAY and the roster. Host-specific team wiring registers the services in teamServicesRegistry and attaches team identity to the context.
| Component | Responsibility |
|---|---|
team.tool.ts | Validate model calls and translate member names into internal identities. |
TeamCoordinator | Persist teams, assignments, messages, deliveries, and dependency transitions. |
TeamExecutionService | Launch independent runAgentLoop member runs, manage ownership, and reconcile exits. |
team-wake-router.ts | Translate events into Lead/member wake requests and UI state updates. |
team-mailbox.ts | Consume the coordinator inbox into the Lead’s per-step reminders. |
Member execution is independent of TaskOrchestrator. It uses kind: "member", TEAM_MEMBER_INSTRUCTIONS, and the member’s specialization instructions. Lead and members share the workspace, while their runtime contexts and model histories are separate. Briefs should assign distinct output paths to avoid competing writes.
The five team tools are currently exposed to both roles. team_create, team_replan, and team_dissolve reject member calls. Member loadouts exclude nested delegation and direct human-interaction tools; questions for the user go through the Lead.
Creation schedules only work that can start
The following illustrative call assumes the named agent types are available in the tool roster. Actual briefs must include the relevant input paths and acceptance criteria.
{
"name": "Auth review",
"output_language": "English",
"members": [
{
"name": "Reviewer",
"agentType": "general",
"title": "Review auth boundaries",
"prompt": "Read /workspace/src/auth. Review session expiry and authorization boundaries. Save evidence and findings to /workspace/auth-review.md. Do not edit source files."
},
{
"name": "Verifier",
"agentType": "general",
"title": "Verify findings",
"prompt": "Read the upstream review and referenced source. Verify each finding and save a checked report to /workspace/auth-verified.md. Distinguish confirmed issues from uncertainty.",
"dependsOn": ["Reviewer"]
}
]
}
team_create accepts 1–8 members per creation call, validates names and cycles, and creates one logical task per member. This input limit is not a claim that team_replan.add enforces a lifetime team-size limit.
A ready assignment becomes in_progress; its launch is deferred until the Lead’s turn commits. An assignment with unmet dependencies stays blocked, without starting a model run. When dependencies finish, it becomes claimable; the member runtime promotes it to in_progress without a model-facing claim tool. Starting work snapshots dependency revisions.
The Lead can add known work later with team_replan. A complete speculative workflow is not required at creation.
Events advance work without polling
| Event | Scheduling effect |
|---|---|
| Message | Wake the recipient; a running member can read it on its next step. |
| Task completed | Wake ready downstream members. Wake the Lead only when no downstream is ready and no task remains in progress. |
| Task failed or cancelled | Wake the Lead for recovery or a decision. |
| Team dissolved | Notify the Lead. |
| Task created, claimed, or member status changed | Update state without starting a model solely for the update. |
team_status returns immediately for both roles and is limited to one call per run. It is not a long-poll operation. Concurrent execution may change the underlying state even during the caller’s run.
Tool results may contain teamId, member IDs, or task IDs for UI bookkeeping. Model inputs use names, and no team tool requires the model to pass a teamId. Work status is nested under each member’s work in the status response; member status and task status are not interchangeable.
Messages can guide work or explicitly yield a run
Messages are persisted before their wake event. The wake event is deferred until the sender’s turn ends. A recipient already running may consume the message before that boundary; deferred wake does not imply a transaction over the sender’s whole turn.
Lead and member steps consume their inboxes into reminders retained for the receiving run. Member reminders also carry the current assignment, revision instruction, and upstream summaries and file paths.
A member that needs an answer before proceeding uses:
{
"to": "Coordinator",
"content": "Which deployment region should I verify? Current notes are in /workspace/notes.md.",
"wait_for_reply": true
}
This stops the run at the tool-step boundary without completing or failing the task. Do not issue complete or unrelated work alongside this wait request. Ordinary messages do not stop the sender. The Lead answers via a normal message; an incoming message can start a new member run. Exit reconciliation checks for replies arriving during cleanup so they are not stranded behind the old run’s lock.
The current run uses a local wait marker; clean finalization persists finishReason: "waiting" in the existing execution record. Restart and ordinary wake paths check the latest completed run and unread mail. With no new mail, a committed wait remains asleep; a post-lock check catches messages arriving during that decision. The logical task stays in_progress, so UI state may still show unfinished work as active. A crash before finalization follows ordinary interrupted-run recovery. Full conversation restoration and matching a reply to a specific question are not implemented.
Delivery, revision, and cancellation preserve distinct meanings
Members call complete({ summary, paths }) when the assignment is fulfilled. A concise finding can use paths: []. If any supplied path is invalid, completion is blocked and downstream tasks remain blocked. For valid delivery, the coordinator records the result, refreshes dependencies, and emits task_completed; successful completion ends the member run.
team_replan supports three operations:
add: create new members with unique names and explicit dependencies, including dependencies on already delivered members.revise: reopen a completed task under the same identity and advance its revision. Delivered downstream tasks becomestaleand rerun in dependency order. Downstream work must be completed or cancelled before revision.cancel: cancel unfinished work and dependent branches. Failed roots keep their failure evidence while their unfinished downstream work is cleared. Replacement nodes do not automatically inherit or redirect old edges.
Replan is sequential, not transactional. Name, graph, post-cancellation dependency validation, and simulated revisions in call order run before mutation. Duplicate or overlapping revisions that invalidate a later operation are rejected before writes. Concurrent changes or infrastructure errors can still leave partial changes. Inspect state before retrying.
The Lead inspects and accepts results, closes the team when no further work remains, and gives the final handover. A Lead complete call does not close the team automatically. Stop/Clear paths also explicitly tear down active members.
Six tables separate task history from execution history
The server schema defines team, team_member, team_task, team_task_delivery, team_message, and team_member_message. Task deliveries preserve revision history; member messages record physical runs, triggers, finish reasons, and usage. Repository adapters implement the shared core contract for each host.
In-process active-run tracking, wake locks, run ownership, and heartbeats prevent overlapping ownership and support cleanup. These mechanisms protect execution; they do not evaluate the correctness of the model’s result or merge competing file edits.
On the client, tool output seeds team state. team:state-update carries task/member transitions and team:member-part carries streamed member content. Snapshot hydration rebuilds state when reopening a conversation. UI events are projections of persisted state, not the authority for scheduling.
Verification separates contracts from model quality
The current focused regression suite covers 60 tests across team services and stop conditions, including message body injection, explicit waiting, SDK loop stopping, reply reconciliation, failed-branch cleanup, completed dependencies, and missing-artifact rejection. TypeScript and formatting checks passed for this implementation round.
These are deterministic checks, including a mock model. They do not establish quality or cost across real models. The next evaluation should cover independent work, dependency handoff, mid-task questions, relevant revisions, failure recovery, and simple requests where delegation adds no value. See the coordination model for remaining boundaries.