This is the preserved partial result for this run. Its original utterance evidence is shown below.
Rows
5
Pass
5 (100%)
Fail
0 (0%)
Unverified
0 (0%)
Pass rate
100%
Avg difficulty
—
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace — needs a human look (that's you 👍/👎-ing it). "Pass rate" = pass ÷ (pass + fail) — it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows).
Why the fails happened — comprehension vs execution vs cosmetic
No fails in this run, or failure classification not present on these rows yet.
Handled correctly? — by expected action
Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed to
N
Correct
Wrong
Unverified
▸ LOG — log the entry
4
4 (100%)
0 (0%)
0 (0%)
No errors — all handled correctly.
▸ UPDATE — update the entry
1
1 (100%)
0 (0%)
0 (0%)
No errors — all handled correctly.
Total
5
5 (100%)
0
0
Accuracy by difficulty
Pending A1's per-utterance difficulty score (requested 2026-07-05) — this bar chart lights up once that lands.
Clarification follow-ups — scored separately
Second turn: app asked, we replied — did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.
Cosmetic only
Not yet classified — pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.
System / infra
Not yet classified — pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.
Latency
Avg (time to ready)
0.3s
p90
0.2s
Max
0.4s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers — click any dot to jump to its detail.
0s
1s
2s
5s
Response path — quick (single response) vs async (an ack like "Working on it…" before the real answer).
Quick response
5
Slowest 5 utterances (click to jump to detail):
"Make sure I text Sarah back and also add a to-do to buy milk."0.4s
"Add a task to renew my passport in the category Admin."0.2s
"Add a to-do to call the dentist."0.2s
"Mark call the dentist as done."0.2s
"Remind me to pick up dry cleaning tomorrow."0.2s
Filter — controls the list below
Pass / Fail / Unverified
PASS 5FAIL 0UNVERIFIED 0
Module (intended for)
todos (5)
Utterance sub-type (within module)
5 shown — 5 pass, 0 fail, 0 unverified
Per-utterance detail
PASSshould log the entry"Add a to-do to call the dentist." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Notes write-truth verified: call the dentist.
Why verdict Notes write-truth verified: call the dentist.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
3 · App data rows written created todo_item: call the dentist
4 · UI did navigated to "todos" screen
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:24:48.607Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Remind me to pick up dry cleaning tomorrow." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Notes write-truth verified: pick up dry cleaning.
Why verdict Notes write-truth verified: pick up dry cleaning.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Added to-do: pick up dry cleaning. Priority 3, due Aug 2. Tell me if the title or due date should be different.
3 · App data rows written created todo_item: pick up dry cleaning
4 · UI did navigated to "todos" screen
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:24:59.927Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Add a task to renew my passport in the category Admin." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Notes write-truth verified: renew my passport.
Why verdict Notes write-truth verified: renew my passport.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
3 · App data rows written created todo_item: renew my passport
4 · UI did navigated to "todos" screen
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:25:11.276Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Make sure I text Sarah back and also add a to-do to buy milk." (difficulty —)0.4s
Verdict Expected LOG — should log the entry. PASS: Notes write-truth verified: buy milk, text Sarah back.
Why verdict Notes write-truth verified: buy milk, text Sarah back.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 · TTS saidno speech captured
2 · Card shown Added 2 to-dos: text Sarah back; buy milk.
3 · App data rows written created todo_item: text Sarah back created todo_item: buy milk
4 · UI did navigated to "todos" screen
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:25:22.921Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould update the entry"Mark call the dentist as done." (difficulty —)0.2s
Verdict Expected UPDATE — should update the entry. PASS: Notes UPDATE write-truth verified (changed:todo_items×1).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Marked done: call the dentist.
3 · App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 · UI did navigated to "todos" screen
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:25:34.191Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)