This is the preserved partial result for this run. Its original utterance evidence is shown below.
Rows
3
Pass
3 (100%)
Fail
0 (0%)
Unverified
0 (0%)
Pass rate
100%
Avg difficulty
—
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace — needs a human look (that's you 👍/👎-ing it). "Pass rate" = pass ÷ (pass + fail) — it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows).
Why the fails happened — comprehension vs execution vs cosmetic
No fails in this run, or failure classification not present on these rows yet.
Handled correctly? — by expected action
Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed to
N
Correct
Wrong
Unverified
▸ LOG — log the entry
3
3 (100%)
0 (0%)
0 (0%)
No errors — all handled correctly.
Total
3
3 (100%)
0
0
Accuracy by difficulty
Pending A1's per-utterance difficulty score (requested 2026-07-05) — this bar chart lights up once that lands.
Clarification follow-ups — scored separately
Second turn: app asked, we replied — did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.
Cosmetic only
Not yet classified — pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.
System / infra
Not yet classified — pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.
Latency
Avg (time to ready)
0.3s
p90
0.3s
Max
0.4s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers — click any dot to jump to its detail.
0s
1s
2s
5s
Response path — quick (single response) vs async (an ack like "Working on it…" before the real answer).
Quick response
3
Slowest 3 utterances (click to jump to detail):
"I had an omelet with three eggs, one ounce of cheddar cheese and a tomato"0.4s
"Ate an apple, a banana, and grapes."0.3s
"I had an omelet"0.2s
Filter — controls the list below
Pass / Fail / Unverified
PASS 3FAIL 0UNVERIFIED 0
Module (intended for)
Other (3)
Utterance sub-type (within module)
3 shown — 3 pass, 0 fail, 0 unverified
Per-utterance detail
PASSshould log the entry"I had an omelet with three eggs, one ounce of cheddar cheese and a tomato" (difficulty —)0.4s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Omelet, Tomato, Egg, Cheddar cheese — card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 · TTS saidno speech captured
2 · Card shown Logged Omelet. Includes three eggs, one ounce of cheddar cheese, and a tomato.
3 · App data rows written created food_log_entry: Omelet ×1 (serving) 351 cal · 26.5g P · 6.8g C · 23.9g F created food_log_entry: Egg ×1 (3 eggs) 215 cal · 18.9g P · 1g C · 14.3g F created food_log_entry: Cheddar cheese ×1 (1 oz) 114 cal · 6.5g P · 1g C · 9.4g F created food_log_entry: Tomato ×1 (1 tomato) 22 cal · 1.1g P · 4.8g C · 0.2g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:54:37.768Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from a manually curated nutrition entry.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had an omelet" (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Omelet — card not captured.
Why verdict Logged (write-truth verified): Omelet — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown I logged omelet — let me know if that's not right.
3 · App data rows written created food_log_entry: Omelet ×1 (serving) 351 cal · 26.5g P · 6.8g C · 23.9g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:54:49.102Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from a manually curated nutrition entry.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Ate an apple, a banana, and grapes." (difficulty —)0.3s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Grapes, Banana, Apple — card not captured.
Why verdict Logged (write-truth verified): Grapes, Banana, Apple — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.3s
1 · TTS saidno speech captured
2 · Card shown Logged an apple, a banana, and Grapes. Assumed catalog default servings where you did not say an amount — tell me if that is not right.
3 · App data rows written created food_log_entry: Apple ×1 (1 apple) 95 cal · 0.5g P · 25.5g C · 0.4g F created food_log_entry: Banana ×1 (1 banana) 105 cal · 1.3g P · 27.1g C · 0.4g F created food_log_entry: Grapes ×1 (100 g) 69 cal · 0.7g P · 18.1g C · 0.2g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:55:00.552Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)