North-star Metric

A north-star metric is not the only important metric. It is the primary metric a team chooses to optimize during a stage of the product. It should represent the direction of value creation and stay close enough to behavior the team can influence.

For an agent product, start with two questions:

  1. Why does the customer pay?
  2. Which kind of successful task best represents that value?

Without those answers, the north-star often collapses into task starts, messages, or MAU.

Selection Principles

A good agent north-star usually has four properties:

PropertyMeaning
Tied to successful outcomesCounts completed work, not just requests
Influenceable by the teamProduct, model, tool, or operations changes can move it
Aligned with the business modelSubscription, usage-based, value-based, and internal tools optimize different things
Protected by guardrailsCannot grow by sacrificing quality, cost, trust, or risk

The easiest mistakes are fast-growing metrics with no quality meaning, such as task starts, message count, token consumption, or session length.

Choose By Business Model

Subscription SaaS

Subscription businesses care about retention and expansion. The primary metric should prove that users continue to delegate real work to the agent, not merely open the product.

Candidate north-star: Weekly Successful Task Users

Definition: users who completed at least one key task this week. “Key task” should be limited to the product’s promised core workflow, not any small action.

Guardrails:

  • Success rate must not fall.
  • Cost per successful task must not run away.
  • Rework rate must not rise.

MAU is not a good north-star here. It is useful for market and retention context, but cannot distinguish “opened the product” from “completed work.”

Usage-Based Pricing

Usage-based businesses care about billable successful outcomes. The primary metric should stay close to revenue without rewarding failed or low-quality completions.

Candidate north-star: Monthly Successful Tasks

Definition: tasks in the month that meet the contractual success criteria and can be billed or delivered.

Guardrails:

  • Failure rate, refund rate, and complaint rate.
  • HITL rate and review minutes.
  • Gross margin per successful task.

If the team optimizes only task count, it may lower pass criteria and make more tasks appear complete. The definition of success must be auditable.

Value-Based / ROI-Anchored Pricing

Value-based businesses care about customer-verifiable value. The primary metric should match the customer’s value model.

Candidate north-star: Verified Value Saved

This can mean hours saved, processing cost reduced, revenue increased, or backlog reduced, but it needs a customer-accepted baseline. Uncalibrated “hours saved” is easy to overstate.

Guardrails:

  • Customer confirmation rate.
  • Coverage of the value calculation.
  • Net value after failure adjustment.

Internal Productivity Tools

Internal tools do not directly monetize usage. The goal is to help employees complete work more efficiently. Avoid inflated hours-saved claims.

Candidate north-star: Successful Key Tasks per Employee

Definition: key agent tasks completed per target employee per month. Key tasks should come from explicit workflows such as report generation, code migration, support handling, or data cleanup.

Guardrails:

  • Employee satisfaction or repeat use.
  • Rework rate.
  • Review minutes.
  • Whether the replaced workflow’s cycle time actually falls.

Required Guardrails

Every north-star needs at least four guardrail categories:

GuardrailPrevents
QualityInflating volume with low-quality tasks
CostBuying growth with loss-making tasks
TrustHiding errors, misreporting success, or eroding user confidence
RiskAutomating irreversible or high-permission actions without control

The north-star points the team in a direction. Guardrails keep it from taking shortcuts.

Diagnostic Metrics

When the north-star moves, diagnostics explain why:

north_star
  = target_users
  × task_start_rate
  × success_rate
  × repeat_use_rate
  × value_per_task

The exact decomposition can vary by product. The important part is knowing whether growth came from more users, deeper penetration, higher success rate, or more valuable work.

Cross-section Connections

Was this page helpful?