what your agent gets asked, where it stops working, and who it happens to.

connect the agent once. every conversation comes back read: what the person wanted, whether they got it, and what got in the way.

works with langchainllamaindexopenai assistantscrewaianthropicany webhook

product questions 2,400 attempts 64% finished

one conversation

t1

summarise what my team shipped.

t7

6 of the 9 documents came back empty. summarising the other 3.

t9

this is missing half of it.

8 further turns, read the same way

what comes back

topic product questions
task summarise what my team shipped failed
intent read it stopped at t7
signal 6 of 9 documents came back empty tone drop at t9
user asked for this more than once
cohort enterprise emea 52% finish this task
task failed on read it 3 of 9 documents used

every span passed.
the task still failed.

the trace says the agent ran. it cannot say six of the nine documents came back empty.

A successful agent run above the result the person received The trace reports every span successful in 1.84 seconds. Below it, of nine documents, three came back with content and six came back empty. agent trace 1.84s every span ok 200 what came back 3 documents read 6 empty

turn a trace into data you can compare.

one trace explains one run. it says nothing about the next four thousand.

abracadabra reads every conversation the same way. the topic, the task, the intents inside it, and the signal on the turn where things changed.

read it failed at t7. that is now a row you can filter and count. it still opens to the turn it came from.

one trace

47 turns · 8.42s
t01 user.message
t07 tool.fetch_docs
t09 user.message
+44 turns
read the same way, every time
topic product questions
task summarise what my team shipped
intent read it
signal 6 of 9 documents empty at t7

one conversation explains what happened.
all of them show what to change.

one broken step and one unlucky run look the same in a single trace. read every attempt and the difference is obvious. this task has three intents. most attempts stop on the second one.

summarise what my team shipped 2,400 attempts

find the source

923.8% stop

read it

52822.9% stop

write the summary

24413.7% stop

finished the task

1,53664%

each bar is the attempts still going at that intent. the coloured end is where they stopped.

864 attempts stopped. 528 of them stopped on the same intent. halve that one and 264 more attempts finish, so 64% becomes 75%.

the same task, finished by cohort

64% baseline

enterprise emea

52%12 under

emea

58%6 under

returning users

61%3 under

enterprise

66%2 over

self-serve

71%7 over

nineteen points from the worst group to the best.

the same task, different results.

64% is the number you would report. nobody actually sits on it.

self-serve finishes this task 71% of the time. enterprise emea finishes it 52% of the time. that is the same agent and the same task, nineteen points apart.

cohorts are built from fields you already send: plan, region, how often somebody comes back. 986 users sit in the worst one.

every number here has a previous one.

new conversations arrive and the same read runs again. so each number has a direction and a size of move.

teal is better, red is worse. set a monitor on a number and you hear about it when it moves, instead of going to look.

against the previous period fastest-rising intent: find the source, +34%

conversations read

12,400

+18%

task success

64%

+6 points

attempts stopped on read it

528

−41%

turns carrying a dissatisfaction signal

3.2%

−1.4 points

monitor

tell me if attempts stopping on read it go back above 600.

watching · 528 now

12,400 conversations, narrowed to one sentence.

each step filters what the step above it left. the last step is a turn somebody typed.

every conversation read 12,400
attempts at this one task 2,400
of those, stopped on read it 528

“this is missing half of it.”

turn 9, in one of the 528

connect the agent. find out how people actually use it.

topics, tasks, intents, signals, people and cohorts are already in the conversations your agent is having. abracadabra is what reads them.

free while in early access, up to 1,000 classified conversations.