Cost Model

An agent cost model should start with formulas, then fill in prices. Model rates, cache multipliers, regional multipliers, batch discounts, and runtime costs all change. The formula is the stable part; price snapshots are inputs.

Basic Formula

The direct cost of one task can be written as:

task_cost =
  input_tokens × input_price
+ cache_write_tokens × cache_write_price
+ cache_read_tokens × cache_read_price
+ output_tokens × output_price
+ tool_runtime_cost
+ external_api_cost
+ retry_cost

If the task requires human review or fallback, add:

total_cost =
  task_cost
+ human_review_minutes × loaded_hourly_cost / 60
+ failure_rate × fallback_cost

The point is to avoid treating the API bill as total cost. The API bill is usually only the most visible layer.

Example Task

The following synthetic example illustrates bill structure.

Task: the user asks an agent to triage the latest 20 emails by project. The agent completes the job in 12 steps: first reading metadata, then applying labels through tools.

Twelve steps is useful because it is a medium-length task. Very short tasks are dominated by startup overhead; long-running tasks are dominated by history and recovery. Medium tasks make it easier to see how the pieces add up.

Reference Bill

Cost Composition — Synthetic 12-Step Task Price snapshot · $0.194 example total · history is the largest category here System+tools Conversation history Tool results Model output 13% 41% 22% 23% <1% Sources: Conversation history = prior turns re-sent as input; Tool results = function returns (enter history); System+tools = static prefix (cached after first step); runtime/storage varies by workload.

The numbers below are synthetic, not a quote. They use one public pricing snapshot and estimated token usage. The absolute values will change; the distribution is here to show which categories can grow.

SourceTokens / usagePricing categoryExample costShare
System prompt (cache hits)2.5K write + 27.5K readscache write / read$0.0189%
Tool descriptions (cache hits)1.2K write + 13.2K readscache write / read$0.0084%
Conversation history (uncached)26.4Kinput$0.07941%
Tool results (uncached)14.4Kinput$0.04322%
Model output3.0Koutput$0.04523%
Sandbox + storage1 session, small storageamortized runtime~$0.001<1%
Total——$0.194100%

Here, cache refers to prompt caching: after an identical static prefix is written once, later requests can read it at a lower price. Prices vary by provider, model, cache TTL, region, and service tier. Use the current official pricing page for production calculations.

Observation 1: History Compounds

In a multi-step agent, each step often needs some history: the user’s goal, previous actions, tool results, errors, and current state. If the full history is replayed, step N carries the content of steps 1 through N-1.

That makes accumulated input grow quickly:

history_tokens ≈ per_step_history × (1 + 2 + ... + n)

Real systems do not have to be strictly O(n²), because they can trim, compact, externalize state, or retrieve only relevant events. The trend still matters: the longer the task, the more important history management becomes.

That is why compaction, memory-system, progress files, structured state, and recoverable sessions are part of the economics, not just user experience.

Observation 2: Caching Depends On Stable Prefixes

Prompt caching works when identical content is read repeatedly. System prompts, tool descriptions, policy blocks, and fixed output formats should stay as stable as possible.

Common ways to break caching:

  • Put per-request user data into the system prompt.
  • Reorder tool descriptions dynamically on every turn.
  • Place timestamps, random IDs, or temporary state inside the static prefix.
  • Constantly rewrite a long prompt to tweak a few words.

Making most of a 12K-token prefix cacheable is often worth more than manually shrinking it to 11K.

Observation 3: Output Is Expensive, But Do Not Just Mute It

For many models, output tokens cost more than input tokens. This encourages fewer long explanations, less unnecessary planning prose, and more direct tool use.

But “shorter output” should not become silence. Useful plans, checkpoints, and error explanations can prevent retries. The better target is:

  • Reduce long reasoning text that is invisible, unverifiable, or not reusable.
  • Make tools return structured results instead of large raw blobs for the model to paraphrase.
  • Emit concise state at key points so users and later agents can resume.

Model Choice Is Multiplicative

Model Choice Scales The Bill 12-step inbox triage · synthetic example · use current pricing for production math Haiku 4.5 ~$0.065 ~3× cheaper than Sonnet Sonnet 4.6 $0.194 baseline · default for most tasks Opus 4.7 ~$0.323 ~1.7× Sonnet · ~5× Haiku — reserve for harder tasks Estimates assume identical token usage. Real cost also depends on step count, retries, cache hits, and pricing changes.

With identical token usage, model selection scales the entire bill by the model’s unit prices. The table below is conceptual and should not be treated as a long-lived price table:

Model tierGood fitCost intuition
Low-cost modelClassification, extraction, formatting, low-risk routine actionsCheap per step, but may need more fallback
Mid-tier modelMost default product agent tasksBalanced cost and reliability
High-capability modelLong-horizon planning, complex code, subjective quality judgment, high-risk decision supportMore expensive per step, but may avoid wrong paths

Model routing is not about always choosing the cheapest model or defaulting to the strongest one. It is about matching task risk, failure cost, and model capability.

Fields To Measure

In production, record these fields rather than only watching the total bill:

FieldWhy it matters
input_tokens / output_tokensShows whether cost comes from input growth or verbose output
cache_write_tokens / cache_read_tokensShows whether static prefixes actually hit cache
tool_result_tokensShows whether tools return too much context
steps / retriesShows wrong paths or loops
model / route_reasonShows whether routing decisions are justified
runtime_secondsShows whether sandbox, browser, or external services matter
human_review_minutesShows whether automation is shifting cost to people
failure_reasonShows whether runaway cost comes from model, tool, permission, data, or product design

These fields feed the budgets, monitoring, and value model in Controls and ROI.

Sources

Was this page helpful?