everything the reading notices, turn by turn.

the topic, the task someone came to finish, the intents inside it, and the signal on each turn.

one conversation, read completely. then the same reading on every other one.

one conversation · product questions · enterprise, EMEA

nobody typed any of the labels on the left and right. they were attached while the conversation was still running.

one conversation shows you something.
the same reading, on all of them, settles an argument.

one reading, at three distances. the turn, the pattern that repeats across thousands of turns, and the person on the other end of both.

a conversation rarely goes wrong all at once.

it goes wrong on one turn. every turn carries its own score, so you get the sentence where it turned and the reply that turned it.

t1 · +0.6
t2
t3
t4
t5
t6
t7 · the turn
t8
t9 · −0.81

“this is missing half of it.”

turn 9 · two turns after 6 of the 9 documents came back empty.

  • scored on the person's turns, so the agent's tone never flatters the number
  • the turn where it changed is marked, so there is something to open
  • one sour conversation is noise. the same drop on the same intent, across hundreds, is a fix

a task does not fail. one step inside it does.

task progress ties every request to the agent action that followed it. the task below has three steps. two of them finished, and the summary still came back half empty.

task · summarise what my team shipped9 turns · 9 documents
complete find the source the engineering space, settled on turn 5 after one correction.
failed read it 6 of the 9 documents came back empty, and the agent moved on.
partial write the summary written from the 3 documents that came back. nothing from the other 6.
task status · failed. the agent said finished. the summary was not.

the same person, every conversation, in one line.

a person's page joins every session they have had. what they attempted, what finished, and which signal keeps turning up. open any point to read the conversation behind it.

session 1

summarise a thread · done

session 3

find the source · done

session 5

shipped summary · partial

session 7

shipped summary · failed

open replay

session 9

same task · failed again

three attempts at one task, and the same empty documents each time. read as three conversations, that is three shrugs. read as one history, it is one thing to fix.

cohorts are built from fields you already send.

plan, region, how often they come back, and any custom field the agent already attaches. combine a few and the group exists.

  • saved once, then available to search, charts and monitors
  • open a cohort to reach its people and their conversations
  • behaviour works the same way: returning users is a field like any other

your agent is not failing everywhere. it is failing here.

conversations group into topics on arrival. each topic breaks into the tasks people attempt inside it. rank the topics by how often those tasks fail, and the argument about what to fix next is mostly over.

36%
29%
22%
18%
14%
9%
product
questions
writing
and editing
finding
a document
account
settings
integrations
billing

task failure rate by topic · 12,400 conversations read, up 18% on the previous period · illustrative data

pick one signal. see which groups are carrying it.

a product-wide rate is an average of groups having very different experiences. plotted against that average, the same signal becomes a short list of who to fix it for first.

signal · documents came back empty · product-wide rate 11%

enterprise, EMEA +19
EMEA +12
returning users +3
enterprise −7
self-serve −9

points above or below the product-wide rate · illustrative data

any signal task progress, tool events, user corrections, repeated intents, clarification failures, sentiment shifts, off-topic drift, escalation intent, and any custom event the agent already emits.
any slice plan, region, returning users, or any field already attached to the person. saved cohorts work here too.

every claim in the summary links to the turn under it.

the summary names the task, the steps that finished, the one that did not, and the signal that decided it. read it out, then open the turn behind whichever sentence somebody doubts.

  • written for a conversation, a person, a topic or a cohort
  • the status comes from the turns, not from a thumbs-down
  • nothing is claimed that has no turn under it

the conversation · 9 turns and 4 tool calls

  1. t1summarise what my team shipped.
  2. t2which space should i read from?
  3. t3the engineering space.
  4. t4found 9 documents in scope.
  5. t5the engineering space, not the whole workspace.
  6. t6re-read. 9 documents in scope.
  7. t76 of the 9 documents came back empty. summarising the other 3.
  8. t8summary written from 3 documents.
  9. t9this is missing half of it.

the summary

someone asked the agent to summarise what their team shipped. the source was corrected once, then settled. reading it did not finish: 6 of the 9 documents came back empty, and the agent wrote the summary from the other 3 anyway.

  • task statusfailed
  • failing intentread it
  • decisive signaldocuments came back empty, turn 7

864 attempts stopped. this is where.

2,400 people asked for this summary and 1,536 finished. the other 864 stopped on one of three steps. the biggest one is first, and that is the order to fix them in.

task · summarise what my team shipped 864 of 2,400 attempts stopped · illustrative data
  • read it528 stopped · 22.9%the documents come back empty and the agent moves on. halve it and 264 more attempts finish. 64% becomes 75%.
  • write the summary244 stopped · 13.7%the summary gets written. it covers what was read, not what was asked for.
  • find the source92 stopped · 3.8%the agent picks the wrong space, or asks which one and never gets an answer.
open the conversations 864 conversations · 986 users in the worst cohort

read one conversation this way. then read all of them.

every conversation read on arrival, and a way to move between the pattern and the turn.

no email, no call. the playground is the live workspace.