Cost Model
An agent cost model should start with formulas, then fill in prices. Model rates, cache multipliers, regional multipliers, batch discounts, and runtime costs all change. The formula is the stable part; price snapshots are inputs.
Basic Formula
The direct cost of one task can be written as:
task_cost =
input_tokens × input_price
+ cache_write_tokens × cache_write_price
+ cache_read_tokens × cache_read_price
+ output_tokens × output_price
+ tool_runtime_cost
+ external_api_cost
+ retry_cost
If the task requires human review or fallback, add:
total_cost =
task_cost
+ human_review_minutes × loaded_hourly_cost / 60
+ failure_rate × fallback_cost
The point is to avoid treating the API bill as total cost. The API bill is usually only the most visible layer.
Example Task
The following synthetic example illustrates bill structure.
Task: the user asks an agent to triage the latest 20 emails by project. The agent completes the job in 12 steps: first reading metadata, then applying labels through tools.
Twelve steps is useful because it is a medium-length task. Very short tasks are dominated by startup overhead; long-running tasks are dominated by history and recovery. Medium tasks make it easier to see how the pieces add up.
Reference Bill
The numbers below are synthetic, not a quote. They use one public pricing snapshot and estimated token usage. The absolute values will change; the distribution is here to show which categories can grow.
| Source | Tokens / usage | Pricing category | Example cost | Share |
|---|---|---|---|---|
| System prompt (cache hits) | 2.5K write + 27.5K reads | cache write / read | $0.018 | 9% |
| Tool descriptions (cache hits) | 1.2K write + 13.2K reads | cache write / read | $0.008 | 4% |
| Conversation history (uncached) | 26.4K | input | $0.079 | 41% |
| Tool results (uncached) | 14.4K | input | $0.043 | 22% |
| Model output | 3.0K | output | $0.045 | 23% |
| Sandbox + storage | 1 session, small storage | amortized runtime | ~$0.001 | <1% |
| Total | — | — | $0.194 | 100% |
Here, cache refers to prompt caching: after an identical static prefix is written once, later requests can read it at a lower price. Prices vary by provider, model, cache TTL, region, and service tier. Use the current official pricing page for production calculations.
Observation 1: History Compounds
In a multi-step agent, each step often needs some history: the user’s goal, previous actions, tool results, errors, and current state. If the full history is replayed, step N carries the content of steps 1 through N-1.
That makes accumulated input grow quickly:
history_tokens ≈ per_step_history × (1 + 2 + ... + n)
Real systems do not have to be strictly O(n²), because they can trim, compact, externalize state, or retrieve only relevant events. The trend still matters: the longer the task, the more important history management becomes.
That is why compaction, memory-system, progress files, structured state, and recoverable sessions are part of the economics, not just user experience.
Observation 2: Caching Depends On Stable Prefixes
Prompt caching works when identical content is read repeatedly. System prompts, tool descriptions, policy blocks, and fixed output formats should stay as stable as possible.
Common ways to break caching:
- Put per-request user data into the system prompt.
- Reorder tool descriptions dynamically on every turn.
- Place timestamps, random IDs, or temporary state inside the static prefix.
- Constantly rewrite a long prompt to tweak a few words.
Making most of a 12K-token prefix cacheable is often worth more than manually shrinking it to 11K.
Observation 3: Output Is Expensive, But Do Not Just Mute It
For many models, output tokens cost more than input tokens. This encourages fewer long explanations, less unnecessary planning prose, and more direct tool use.
But “shorter output” should not become silence. Useful plans, checkpoints, and error explanations can prevent retries. The better target is:
- Reduce long reasoning text that is invisible, unverifiable, or not reusable.
- Make tools return structured results instead of large raw blobs for the model to paraphrase.
- Emit concise state at key points so users and later agents can resume.
Model Choice Is Multiplicative
With identical token usage, model selection scales the entire bill by the model’s unit prices. The table below is conceptual and should not be treated as a long-lived price table:
| Model tier | Good fit | Cost intuition |
|---|---|---|
| Low-cost model | Classification, extraction, formatting, low-risk routine actions | Cheap per step, but may need more fallback |
| Mid-tier model | Most default product agent tasks | Balanced cost and reliability |
| High-capability model | Long-horizon planning, complex code, subjective quality judgment, high-risk decision support | More expensive per step, but may avoid wrong paths |
Model routing is not about always choosing the cheapest model or defaulting to the strongest one. It is about matching task risk, failure cost, and model capability.
Fields To Measure
In production, record these fields rather than only watching the total bill:
| Field | Why it matters |
|---|---|
input_tokens / output_tokens | Shows whether cost comes from input growth or verbose output |
cache_write_tokens / cache_read_tokens | Shows whether static prefixes actually hit cache |
tool_result_tokens | Shows whether tools return too much context |
steps / retries | Shows wrong paths or loops |
model / route_reason | Shows whether routing decisions are justified |
runtime_seconds | Shows whether sandbox, browser, or external services matter |
human_review_minutes | Shows whether automation is shifting cost to people |
failure_reason | Shows whether runaway cost comes from model, tool, permission, data, or product design |
These fields feed the budgets, monitoring, and value model in Controls and ROI.