Evaluation & Observability
Observed facts and evaluation judgments are not one score
Observability answers what happened during a run. Evaluation asks whether the result met its objective. The former should be faithful and replayable; the latter depends on acceptance criteria and may combine rules, models, and human judgment. aibuddy persists task facts first, then derives views for diagnosis, cost, and quality instead of writing one composite score as system truth.
A successful tool response proves only that an action returned. A natural agent stop proves only that a run ended. File contents, test results, page state, cited evidence, or user confirmation determine whether the task itself is acceptable.
Task outcomes are independent of model finish reasons
A step finishReason explains why one generation stopped, such as stop, tool-calls, length, or error. Task lifecycle state must instead say whether the user’s work can still progress. aibuddy stores isActive, awaitingUserInputAt, and lastFinishReason separately rather than inferring every outcome from one boolean.
The admin projection derives six states:
| State | Basis |
|---|---|
running | The task is active and not waiting for user input |
awaiting_input | The task is active with unresolved HITL interaction |
completed | The task ended naturally and persisted completed |
failed | A model or execution error terminated the task |
aborted | A user or administrator deliberately stopped the task |
interrupted | Shutdown or another system event interrupted the run |
An HITL pause is not counted as completion, and deployment shutdown does not contaminate user-abort metrics. Only legacy rows without lastFinishReason use the compatibility mapping to completed.
Step telemetry pairs execution plans with actual results
task_step_info stores one process-metric row per model step. The prepareStep side records predicted input, visible tool count, compaction landing, and reminder activity. The onStepEnd side adds actual tokens, cache reads and writes, reasoning tokens, tool calls, model, duration, warnings, and finish reason. Both sides meet in one step record.
This pairing answers concrete questions: how far estimated context differed from actual input, which step compacted, how much tool output expanded the following step, whether cache reuse occurred, and where abnormal cost began. Step records keep metrics and tool names without duplicating full prompts or model responses.
Task budget uses a separate append-on-change snapshot. A new row is added only when the context window, system instructions, tool schemas, model, or effective loadout changes. Stable runs do not repeat the same configuration on every step. Admin views load steps, subagents, loadout, and compaction snapshots lazily so a long task does not put its full trace in the base detail response.
Model metering and operational tracing use separate lanes
Every model entry eventually crosses the @aibuddy/ai call boundary or the Agent Loop step hook and emits a common ModelInteraction. The event carries source attribution, user and task context, provider, model, BYOK state, token detail, outcome, and duration. UsageSource distinguishes task, chat, team, subagent, memory, compaction, knowledge, eval, and other sources. Missing attribution is still recorded as unknown and emits a warning rather than disappearing silently.
The server usage sink asynchronously appends these events to model_usage. A sink failure is logged without breaking the model path, and the table deliberately avoids foreign keys to task or user rows that may later be pruned. Metering therefore remains visible even when a separate billing policy decides that an interaction consumes no credits.
Operational tracing serves a different query. The server can emit model and agent spans through OpenTelemetry and correlate them with logs using one trace ID; Desktop does not register this integration. aiTelemetry disables input and output capture by default so complete prompts and responses do not enter the trace backend. Content capture requires an explicitly configured diagnostic process with OTEL_CAPTURE_CONTENT=true.
| Lane | Primary question | Default content boundary |
|---|---|---|
| Task records | What did the user receive, and how did the task end? | Messages, artifacts, state, and necessary metadata |
| Step telemetry | Which step had context, tool, or cost anomalies? | Metrics without full prompts |
| Model usage | Which capability consumed which model resources? | Attribution, model, tokens, outcome, and duration |
| Logs and traces | Where did a cross-process call slow down or fail? | Structured events; model content disabled by default |
The same evidence supports different inspection views
The admin task evaluation view reads real task state and step telemetry to derive completion, budget, efficiency, control, evidence, and risky-step views. Its current quality scores are deterministic heuristics over status, tool counts, finish reasons, tokens, and compaction data. They are useful for triage, but they are neither LLM-as-judge results nor user acceptance.
User feedback uses a more direct unit. Each assistant message can store up, down, or no rating. Because feedback belongs to one response rather than an entire user or agent, it can be joined back to the task state, steps, and model usage that produced it. It still expresses user preference and cannot alone establish factual correctness or safety.
The aggregate Evaluations page currently uses mock data to demonstrate target views for success rate, version comparison, consistency, tool accuracy, and takeover rate. It is not connected to one production evaluation repository, so those figures must not be presented as measured online performance.
Compaction evaluation has a closed targeted-judge path
Context compaction is a narrower area with a real judge path. An administrator selects a compacted part or rolling summary. The system reconstructs the original from persisted task messages, reads the compacted form from the durable snapshot, and sends both to an independent judge.
The structured result contains:
| Field | Meaning |
|---|---|
verdict | ok, minor, or major information-loss severity |
faithful | Whether the compacted form distorts or fabricates information |
lostItems | Important facts that were lost or obscured |
reasoning | A concise basis for the verdict |
promptFindings | Minimal attributable instruction changes for summary generation only |
The judge also distinguishes content recoverable through view_tool_call from content with no recovery handle, avoiding a severe-loss verdict merely because bulky recoverable output was removed. Oversized inputs retain their head and tail, and the judgment is explicitly limited to visible content.
The verdict streams as structured output. For an inactive task, it persists in a separate eval column beside the compaction snapshot. For an active task, it remains session-only: the runtime may compact the same key again, and an asynchronous old verdict must not be written against new content. This restriction keeps the evaluated object aligned with the stored judgment.
Attribute the failure layer before choosing a fix
| Failure layer | Available evidence | First repair surface |
|---|---|---|
| Objective and acceptance | Final artifact misses the user’s criteria | Task contract, acceptance rule, or human gate |
| Model judgment | Wrong plan or conclusion | Instructions, examples, model, or targeted eval |
| Context | Missing constraint, failed retrieval, or lossy compaction | Assembly, memory, knowledge, or compaction |
| Tool execution | Invalid parameter, denial, timeout, or partial side effect | Schema, host, authorization, or recovery path |
| Runtime platform | Queue, connection, storage, lock, or deployment interruption | Platform state, logs, and cross-service trace |
| Resource efficiency | Correct result with abnormal steps, tokens, or latency | Loadout, cache, scheduling, or stop conditions |
Without this attribution, teams tend to treat every failure as a prompt problem. Observability is valuable not because it stores many fields, but because it locates the system layer that should change.
Current capability boundary
aibuddy has durable task outcomes, per-step process telemetry, budget and loadout snapshots, unified model metering, message-level feedback, OpenTelemetry integration, and on-demand compaction judging. It does not yet have one evaluation backend that joins online samples, user feedback, human labels, automated judges, and versioned eval sets.
The system can therefore answer how a task ran, where it became abnormal, and how many resources it consumed, while performing targeted quality audits of compaction. System-wide success rate, cross-version quality gain, and consistency require a repeatable dataset and a unified verdict policy before they can govern releases.
Implementation anchors
| Responsibility | Module |
|---|---|
| Task outcome and HITL classification | task-service.finalizeTurn |
| Per-step process metrics | task_step_info / createOnStepEnd |
| Budget and loadout snapshots | task_budget / saveBudgetSnapshot |
| Common model events and source attribution | ModelInteraction / UsageSource |
| Product usage persistence | usage-sink / model_usage |
| Operational trace content boundary | aiTelemetry / registerTelemetry |
| Compaction audit and structured judge | streamCompactionEval / compactionEvalSchema |
Related reading
- Observability Stack for deployment boundaries between logs, metrics, and traces.
- Dashboards for deriving metrics from operational questions.
- Datasets and Metrics for repeatable cases and success criteria.
- Continuous Evolution for constraining changes with evaluation evidence.