Agent Economics
Traditional SaaS economics often starts with “users × monthly price” on the revenue side and servers, bandwidth, and storage on the cost side. Agent products do not fit that shape. Users consume inference, tool calls, repeated context, retries, and human handoffs rather than just seats.
The goal of agent economics is not to memorize today’s model prices. It is to answer three durable questions:
- Why does one task cost what it costs?
- Which costs grow with task length?
- How do cost controls and business value fit into the same model?
Where Cost Comes From
One agent task usually contains six cost categories:
| Cost | Meaning | Common controls |
|---|---|---|
| Input tokens | System prompt, user input, history, and tool results entering context | Prompt caching, trimming, retrieval granularity |
| Cached tokens | Static prefixes read from cache at lower prices | Keep system prompts and tool descriptions stable |
| Output tokens | Natural language, plans, explanations, or tool arguments generated by the model | Reduce unnecessary reasoning output; prefer direct tool calls |
| Tool / runtime | Browser, sandbox, database, external API, file system, or session infrastructure | Tool limits, batching, sandbox lifecycle management |
| Retry / failure | Extra inference, rollback, and redo work after wrong paths | Evals, budget stops, user confirmation points |
| Human review | Review, takeover, correction, and process training | Place HITL only at high-risk points |
Tokens are the easiest part to measure, but they are not the whole cost. A seemingly cheap agent can still have poor economics if it fails often and requires humans to clean up the work.
Three Dimensions Of A Token
Each token consumes three resources at once:
| Resource | Meaning | Business impact |
|---|---|---|
| Money | Token count × current model price | Bill and gross margin |
| Latency | Input processing and output generation take time | User-facing wait time |
| Capacity | Tokens occupy the context window | Task depth and retained state |
Before optimizing, name the target. Many choices are trade-offs rather than pure improvements:
- Compacting history can reduce future input, but compaction itself costs tokens and may lose detail.
- A stronger model costs more per token, but may take fewer steps and retry less.
- A cheaper model lowers per-step cost, but may increase step count, failure rate, and human takeover.
- A longer context can reduce handoff overhead, but may also make every step carry more history.
Agent cost work is therefore not just “save tokens.” It is deciding which tokens are worth spending and which tasks should not be delegated to an agent.
Three Governing Principles
Keep static content cacheable. If the system prompt, tool descriptions, or policy blocks change on every step, prompt caching breaks. Preserving a large static prefix usually matters more than manually deleting a few hundred tokens.
Long-running history compounds. Multi-step agents often resend prior messages and tool results into context. Without trimming, compaction, or externalized state, history cost grows quickly with step count.
Failure is rarely cheaper than success. Tokens, tool calls, and human time spent on a failed path are sunk. A late-stage failure in a long task can cost more to repair than one successful run.
A Reference Bill
The following example is for intuition, not a quote. Suppose a user asks an agent to triage the latest 20 emails by project, and the task runs for 12 steps. Using one public pricing snapshot and synthetic token usage, the bill is approximately $0.19.
The rough distribution is:
- 41% Conversation history: prior messages and tool results re-enter later inputs.
- 23% Model output: responses, plans, and tool arguments.
- 22% Tool results: returned tool content that later becomes context.
- 13% System prompt + tool descriptions: static prefix, assuming strong cache hits.
- < 1% Sandbox + storage: small in this example, but workload-dependent in real systems.
This distribution is not a law. It shows that in medium-length, multi-step agent tasks with history replay, history, tool results, and output often deserve more attention than startup prompt overhead.
Section Guide
| Section | What it covers | Audience |
|---|---|---|
| Cost Model | Formulas and a synthetic example for input, cache, output, tools, failures, and human review | Engineering, product, finance |
| Controls and ROI | Budgets, model routing, caching, circuit breakers, monitoring, and conservative cost-benefit estimation | Operations, sales, procurement |
If you read only one page, start with Cost Model. It turns “agents are expensive” into fields that can be measured.