what your agent gets asked, where it stops working, and who it happens to.
connect the agent once. every conversation comes back read: what the person wanted, whether they got it, and what got in the way.
works with langchainllamaindexopenai assistantscrewaianthropicany webhook
one conversation
summarise what my team shipped.
6 of the 9 documents came back empty. summarising the other 3.
this is missing half of it.
8 further turns, read the same way
what comes back
every span passed.
the task still failed.
the trace says the agent ran. it cannot say six of the nine documents came back empty.
turn a trace into data you can compare.
one trace explains one run. it says nothing about the next four thousand.
abracadabra reads every conversation the same way. the topic, the task, the intents inside it, and the signal on the turn where things changed.
read it failed at t7. that is now a row you can filter and count. it still opens to the turn it came from.
one trace
47 turns · 8.42sone conversation explains what happened.
all of them show what to change.
one broken step and one unlucky run look the same in a single trace. read every attempt and the difference is obvious. this task has three intents. most attempts stop on the second one.
find the source
923.8% stop
read it
52822.9% stop
write the summary
24413.7% stop
finished the task
1,53664%
864 attempts stopped. 528 of them stopped on the same intent. halve that one and 264 more attempts finish, so 64% becomes 75%.
the same task, finished by cohort
enterprise emea
52%12 under
emea
58%6 under
returning users
61%3 under
enterprise
66%2 over
self-serve
71%7 over
the same task, different results.
64% is the number you would report. nobody actually sits on it.
self-serve finishes this task 71% of the time. enterprise emea finishes it 52% of the time. that is the same agent and the same task, nineteen points apart.
cohorts are built from fields you already send: plan, region, how often somebody comes back. 986 users sit in the worst one.
every number here has a previous one.
new conversations arrive and the same read runs again. so each number has a direction and a size of move.
teal is better, red is worse. set a monitor on a number and you hear about it when it moves, instead of going to look.
conversations read
12,400
+18%
task success
64%
+6 points
attempts stopped on read it
528
−41%
turns carrying a dissatisfaction signal
3.2%
−1.4 points
tell me if attempts stopping on read it go back above 600.
watching · 528 now12,400 conversations, narrowed to one sentence.
each step filters what the step above it left. the last step is a turn somebody typed.
“this is missing half of it.”
connect the agent. find out how people actually use it.
topics, tasks, intents, signals, people and cohorts are already in the conversations your agent is having. abracadabra is what reads them.
free while in early access, up to 1,000 classified conversations.