Governance and Security (G)
This page corresponds to §9 of Agent Harness Engineering: A Survey. Governance and Security is the second layer the taxonomy promotes to first-class status. It asks how agent behavior is constrained, made safer, and made accountable.
Agents now execute shell commands, send email, submit code, browse sites, and call third-party APIs. Production systems therefore need answers to two questions: under what constraints may the agent act, and who is accountable when those constraints fail?
Permission Models And Identity
Access control for agents is harder than traditional RBAC/ABAC because the needed tool set often depends on a natural language task that was not known at deployment time.
| Granularity | Mechanism | Examples |
|---|---|---|
| Static permission boundary | Fixed allow/deny lists and sandbox limits | Codex restricted sandboxes, Gemini CLI allow/deny |
| Context-dependent privilege control | Evaluate predicates over tool, arguments, environment, and task context | Progent, Conseca |
| Agent identity and delegation | Bind user intent to agent authority and audit chain | Authenticated Delegation, SAGA, IsolateGPT |
| Credential management | Keep secrets in a vault; expose placeholders to the model | Skyvern |
| Web-level permission coordination | Site-declared action permissions and rate/confirmation requirements | agent-permissions.json |
Open problems include portable policy languages, task-scoped identities, and credential lifecycles for long-running sessions.
Lifecycle Hooks
Permission models define what is allowed. Lifecycle hooks define when checks fire.
| Hook | Location | Purpose | Examples |
|---|---|---|---|
| H1 input guardrail | Before the LLM | Detect injection payloads in user input or retrieved content | PromptShield, DataSentinel |
| H2 action validation | Before tool execution | Check proposed actions and multi-agent control-flow transitions | ShieldAgent, ControlValve |
| H3 information flow control | After tool execution | Track provenance and prevent untrusted data from influencing control flow | CaMeL |
| H4 human-in-the-loop | Before consequential actions or at handoff points | Gate destructive or out-of-scope actions on approval | Codex, Gemini CLI, Cursor, OpenHands |
Human approval is not free. Prior permission-UI research shows that users often ignore or misunderstand dialogs; agent approval prompts face similar habituation risks. Frequent prompts teach reflexive approval, while sparse prompts leave coverage gaps.
Hook APIs are heterogeneous, and stacked hooks can interact badly. An upstream sanitizer might remove the signal a downstream detector depends on.
Component Hardening
Hooks enforce policy around the loop. Component hardening makes the model and tools less vulnerable before hooks fire.
- Model hardening: instruction hierarchy and SecAlign train models to prioritize privileged instructions over untrusted data.
- Classifier runtime hardening: Llama Guard-style models screen input/output with configurable taxonomies.
- Tool and MCP security: ETDI signs and versions tool definitions; SAFEFLOW adds protocol-level information flow control and transactional rollback.
- Supply chain hardening: package hallucination and slopsquatting show that agents can introduce compromised dependencies even when the immediate tool call looks valid.
No single hardening layer covers all threats. Hardened models can still misuse compromised tools; signed tools do not prevent a jailbroken model from choosing unsafe actions.
Declarative Constitutions
Governance rules become more inspectable when externalized:
| Level | Mechanism | Signal |
|---|---|---|
| Training-time constitution | Shape model alignment and preference hierarchy | Powerful but hard to inspect and update. |
| Deployment-time YAML | Versioned, diffable rules for pipeline mode, risk mode, allow/deny, budgets, and audit destinations | Directly inspectable by non-model engineers. |
| Programmable policy language | Predicates, quantifiers, automata, or UI transition DSLs | More formal power, less portability. |
The open question is composition: can deployment-time governance reliably override training-time dispositions, and can policy schemas become portable across harnesses?
Audit Infrastructure
Governance requires accountability. A replayable audit record should include trace IDs, principal identity, tool calls, policy decisions and versions, execution results, resource costs, and integrity hashes for relevant inputs/outputs.
Most systems record only a subset. Few sign or hash enough of the record to protect it from a compromised process.
The paper distinguishes:
- per-action monitoring: cheap and local, but misses slow multi-step attacks;
- trajectory-level audit: better for multi-step patterns, but higher latency and harder triggering.
Practical systems often run per-action checks inline and trajectory-level analysis asynchronously over audit logs.
Governance Coverage Is Sparse
The survey’s governance coverage matrix encodes representative systems across permissions, hooks, hardening, constitutions, audit, and multi-agent governance. The important result is that no system covers the whole surface.
| System family | Strength | Gap |
|---|---|---|
| Codex / Gemini CLI / OpenHands | Sandboxes, action approval, partial audit | Fine-grained identity and formal policy portability. |
| AutoHarness | Declarative risk-tiered governance | Schema portability and external interoperability. |
| Progent / Conseca | Context-dependent permissions and generated policy checks | General runtime enforcement and standardization. |
| CaMeL / SAFEFLOW | Information flow and transaction semantics | Integration cost and flexibility tradeoffs. |
| SAGA / Authenticated Delegation | Identity, delegation, short-lived credentials | End-to-end integration with tools, traces, and audit. |
| AgentSpec / AgentDoG | Programmable hooks and trajectory diagnosis | Hook standards and compositional effects. |
Governance is therefore not a security toggle. It is a cross-layer control plane spanning identity, tools, context, lifecycle, audit, and human authority.
Security Landscape
The paper maps governance mechanisms to risk categories from the broader agent-security literature:
- untrusted interfaces;
- wrong instruction following;
- unconstrained data flow;
- hallucination;
- data leakage;
- unauthorized action;
- resource exhaustion.
It also notes contextual security as a proposed extension beyond the CIA triad. That framing is useful, but the paper treats it as an emerging proposal rather than an established standard.
Research Directions
Governance-specific open directions include:
- standardized policy and audit languages;
- formal guarantees over governance pipelines;
- adaptive governance and governance of policy generators;
- long-horizon governance with renewal, revocation, and audit over evolving sessions;
- usable permission UIs, audit dashboards, and constitution editors;
- cross-layer coherence between training-time, deployment-time, and runtime governance;
- end-to-end supply chain governance;
- unified adversarial benchmarks for full governance stacks.