A production agent can return a valid answer, complete every tool call, and stay inside latency and cost targets. None of those facts prove that the customer achieved what they came to do.
The missing layer is the customer job
Traditional analytics are useful, but each system sees a different fragment. Event analytics can tell you that messages were sent. Trace tools can tell you that retrieval and tool execution worked. Transcript review can show a memorable failure. The product question sits across all three: did this population of customers complete the job?
Consider a refund-status agent. Its identity tool succeeds, the model responds, and the session ends without an exception. A customer who already verified may still be asked to verify again. The agent run is operationally healthy; the customer experience is not.
Did the run and its tools complete?
What did the customer and agent actually do?
Who else experienced the same pattern?
Did the customer job improve in the next window?
Separate observation from interpretation
A useful outcome view does not turn every pattern into a causal claim. It labels how far the evidence goes. An exact tool event is observed. A repeated pattern across a cohort is associated. A modeled business impact is estimated. If the evidence is absent, the answer is missing—not zero.
Product truth becomes more useful when a team can see both the claim and the boundary of the claim.
Abracadabra field notes
This distinction changes the quality of the product conversation. Teams can act on a well-supported pattern without pretending they have proven revenue causality. They can also identify the external proof—billing data, a controlled experiment, or human confirmation—needed to make a stronger claim.
Every important number should open into its definition, population, time window, and supporting conversations.
Turn evidence into an operating loop
Once the failure pattern is visible, preserve its scope. Create an issue that owns the problem, define the outcome that should improve, and monitor whether the condition returns. The same cohort should travel through all three steps so the team does not quietly change the denominator halfway through the work.
- Map demand.Group the language customers use into intents, outcomes, signals, and reusable cohorts.
- Inspect evidence.Open affected people, exact turns, tool activity, and representative conversations.
- Own the response.Promote the supported finding into product work without losing its scope or provenance.
- Measure the next window.Compare the same definition over time and keep unresolved conversations available for review.
Start with one outcome your team can change
The first useful outcome does not need to be a universal business metric. Choose a customer job with enough conversation volume, a clear completion definition, and a product owner who can respond. Resolution, completion, activation, and escalation are often better starting points than a broad promise such as “improve satisfaction.”
Then ask a concrete question: which customers are failing, what pattern is shared across their conversations, and what should be different after the next product change? That question gives the team a scope it can inspect, own, and measure.
Put the model into practiceBring one live agent and one outcome you need to improve.Request design-partner access