Evaluation & Observability

Observed facts and evaluation judgments are not one score

Observability answers what happened during a run. Evaluation asks whether the result met its objective. The former should be faithful and replayable; the latter depends on acceptance criteria and may combine rules, models, and human judgment. aibuddy persists task facts first, then derives views for diagnosis, cost, and quality instead of writing one composite score as system truth.

aibuddy evaluation and observability evidence chainA task run persists state, messages, step metrics, and model usage while producing operational traces without model content. Real evidence drives task inspection, compaction judging, and message feedback; a unified system-level evaluation backend is not yet connected. Persist run facts before forming quality judgmentsOne evidence chain supports diagnosis, cost attribution, and evaluation without sharing one conclusion. Task executionAgent Loop · tools · HITL · subagentsA model stop is not task acceptance. Operational logs + tracescross-process timing · diagnosismodel content disabled by default Durable evidenceStores reviewable facts without deciding quality in advance Task state + messagesoutcome · HITL · artifacts Step telemetryplanned · actual · tools Model usagesource · tokens · outcome · time Budget + compactionloadout · source · reduced form not unified Task rule inspectionstate · budget · efficiency · riskdeterministic projection of real data Compaction judgesource ⇄ compacted formstreamed review, safe persistence Message feedbackup · down · nullpreference signal, not factual verdict Unified eval backendonline samples · judge · baselinesaggregate view currently uses mock data Reviewable facts, attributable judgments, and explicit gaps.

A successful tool response proves only that an action returned. A natural agent stop proves only that a run ended. File contents, test results, page state, cited evidence, or user confirmation determine whether the task itself is acceptable.

Task outcomes are independent of model finish reasons

A step finishReason explains why one generation stopped, such as stop, tool-calls, length, or error. Task lifecycle state must instead say whether the user’s work can still progress. aibuddy stores isActive, awaitingUserInputAt, and lastFinishReason separately rather than inferring every outcome from one boolean.

The admin projection derives six states:

StateBasis
runningThe task is active and not waiting for user input
awaiting_inputThe task is active with unresolved HITL interaction
completedThe task ended naturally and persisted completed
failedA model or execution error terminated the task
abortedA user or administrator deliberately stopped the task
interruptedShutdown or another system event interrupted the run

An HITL pause is not counted as completion, and deployment shutdown does not contaminate user-abort metrics. Only legacy rows without lastFinishReason use the compatibility mapping to completed.

Step telemetry pairs execution plans with actual results

task_step_info stores one process-metric row per model step. The prepareStep side records predicted input, visible tool count, compaction landing, and reminder activity. The onStepEnd side adds actual tokens, cache reads and writes, reasoning tokens, tool calls, model, duration, warnings, and finish reason. Both sides meet in one step record.

This pairing answers concrete questions: how far estimated context differed from actual input, which step compacted, how much tool output expanded the following step, whether cache reuse occurred, and where abnormal cost began. Step records keep metrics and tool names without duplicating full prompts or model responses.

Task budget uses a separate append-on-change snapshot. A new row is added only when the context window, system instructions, tool schemas, model, or effective loadout changes. Stable runs do not repeat the same configuration on every step. Admin views load steps, subagents, loadout, and compaction snapshots lazily so a long task does not put its full trace in the base detail response.

Model metering and operational tracing use separate lanes

Every model entry eventually crosses the @aibuddy/ai call boundary or the Agent Loop step hook and emits a common ModelInteraction. The event carries source attribution, user and task context, provider, model, BYOK state, token detail, outcome, and duration. UsageSource distinguishes task, chat, team, subagent, memory, compaction, knowledge, eval, and other sources. Missing attribution is still recorded as unknown and emits a warning rather than disappearing silently.

The server usage sink asynchronously appends these events to model_usage. A sink failure is logged without breaking the model path, and the table deliberately avoids foreign keys to task or user rows that may later be pruned. Metering therefore remains visible even when a separate billing policy decides that an interaction consumes no credits.

Operational tracing serves a different query. The server can emit model and agent spans through OpenTelemetry and correlate them with logs using one trace ID; Desktop does not register this integration. aiTelemetry disables input and output capture by default so complete prompts and responses do not enter the trace backend. Content capture requires an explicitly configured diagnostic process with OTEL_CAPTURE_CONTENT=true.

LanePrimary questionDefault content boundary
Task recordsWhat did the user receive, and how did the task end?Messages, artifacts, state, and necessary metadata
Step telemetryWhich step had context, tool, or cost anomalies?Metrics without full prompts
Model usageWhich capability consumed which model resources?Attribution, model, tokens, outcome, and duration
Logs and tracesWhere did a cross-process call slow down or fail?Structured events; model content disabled by default

The same evidence supports different inspection views

The admin task evaluation view reads real task state and step telemetry to derive completion, budget, efficiency, control, evidence, and risky-step views. Its current quality scores are deterministic heuristics over status, tool counts, finish reasons, tokens, and compaction data. They are useful for triage, but they are neither LLM-as-judge results nor user acceptance.

User feedback uses a more direct unit. Each assistant message can store up, down, or no rating. Because feedback belongs to one response rather than an entire user or agent, it can be joined back to the task state, steps, and model usage that produced it. It still expresses user preference and cannot alone establish factual correctness or safety.

The aggregate Evaluations page currently uses mock data to demonstrate target views for success rate, version comparison, consistency, tool accuracy, and takeover rate. It is not connected to one production evaluation repository, so those figures must not be presented as measured online performance.

Compaction evaluation has a closed targeted-judge path

Context compaction is a narrower area with a real judge path. An administrator selects a compacted part or rolling summary. The system reconstructs the original from persisted task messages, reads the compacted form from the durable snapshot, and sends both to an independent judge.

The structured result contains:

FieldMeaning
verdictok, minor, or major information-loss severity
faithfulWhether the compacted form distorts or fabricates information
lostItemsImportant facts that were lost or obscured
reasoningA concise basis for the verdict
promptFindingsMinimal attributable instruction changes for summary generation only

The judge also distinguishes content recoverable through view_tool_call from content with no recovery handle, avoiding a severe-loss verdict merely because bulky recoverable output was removed. Oversized inputs retain their head and tail, and the judgment is explicitly limited to visible content.

The verdict streams as structured output. For an inactive task, it persists in a separate eval column beside the compaction snapshot. For an active task, it remains session-only: the runtime may compact the same key again, and an asynchronous old verdict must not be written against new content. This restriction keeps the evaluated object aligned with the stored judgment.

Attribute the failure layer before choosing a fix

Failure layerAvailable evidenceFirst repair surface
Objective and acceptanceFinal artifact misses the user’s criteriaTask contract, acceptance rule, or human gate
Model judgmentWrong plan or conclusionInstructions, examples, model, or targeted eval
ContextMissing constraint, failed retrieval, or lossy compactionAssembly, memory, knowledge, or compaction
Tool executionInvalid parameter, denial, timeout, or partial side effectSchema, host, authorization, or recovery path
Runtime platformQueue, connection, storage, lock, or deployment interruptionPlatform state, logs, and cross-service trace
Resource efficiencyCorrect result with abnormal steps, tokens, or latencyLoadout, cache, scheduling, or stop conditions

Without this attribution, teams tend to treat every failure as a prompt problem. Observability is valuable not because it stores many fields, but because it locates the system layer that should change.

Current capability boundary

aibuddy has durable task outcomes, per-step process telemetry, budget and loadout snapshots, unified model metering, message-level feedback, OpenTelemetry integration, and on-demand compaction judging. It does not yet have one evaluation backend that joins online samples, user feedback, human labels, automated judges, and versioned eval sets.

The system can therefore answer how a task ran, where it became abnormal, and how many resources it consumed, while performing targeted quality audits of compaction. System-wide success rate, cross-version quality gain, and consistency require a repeatable dataset and a unified verdict policy before they can govern releases.

Implementation anchors

ResponsibilityModule
Task outcome and HITL classificationtask-service.finalizeTurn
Per-step process metricstask_step_info / createOnStepEnd
Budget and loadout snapshotstask_budget / saveBudgetSnapshot
Common model events and source attributionModelInteraction / UsageSource
Product usage persistenceusage-sink / model_usage
Operational trace content boundaryaiTelemetry / registerTelemetry
Compaction audit and structured judgestreamCompactionEval / compactionEvalSchema
Was this page helpful?