This is the preserved partial result for this run. Its original utterance evidence is shown below.
Rows
12
Pass
11 (92%)
Fail
1 (8%)
Unverified
0 (0%)
Pass rate
92%
Avg difficulty
—
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace — needs a human look (that's you 👍/👎-ing it). "Pass rate" = pass ÷ (pass + fail) — it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows).
Why the fails happened — comprehension vs execution vs cosmetic
Comprehension — picked the wrong action/target (the hard problem)
1 (100%)
resolution — 1 (100% of comprehension)
"I had one cup cooked oatmeal with no toppings." — Over-asked: asked instead of logging (no saved row).
Of 1 fails: if most are comprehension, that's the hard problem (the app didn't figure out the right action/target); execution/cosmetic means the decision was right but something downstream broke. Click a bar for the sub-split + example utterances.
Handled correctly? — by expected action
Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed to
N
Correct
Wrong
Unverified
▸ LOG — log the entry
12
11 (92%)
1 (8%)
0 (0%)
1 handled wrong — click one to jump to its full detail below
"I had one cup cooked oatmeal with no toppings."
OVER-ASK — asked instead of logging (no saved row).
Total
12
11 (92%)
1
0
Accuracy by difficulty
Pending A1's per-utterance difficulty score (requested 2026-07-05) — this bar chart lights up once that lands.
Clarification follow-ups — scored separately
Second turn: app asked, we replied — did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.
Cosmetic only
Not yet classified — pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.
System / infra
Not yet classified — pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.
Latency
Avg (time to ready)
0.7s
p90
0.4s
Max
5.6s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers — click any dot to jump to its detail.
0s
1s
2s
5s
6s
Response path — quick (single response) vs async (an ack like "Working on it…" before the real answer).
Quick response
12
Slowest 8 utterances (click to jump to detail):
"I had one cup cooked oatmeal with no toppings."5.6s
"Track one Quest chocolate chip cookie dough bar."0.5s
"Add two hard boiled eggs for lunch."0.4s
"I drank one cup two percent milk."0.3s
"Log one cup blueberries as a snack."0.3s
"Log one medium banana for the sequence."0.2s
"Record four ounces grilled chicken."0.2s
"Put down four ounces grilled chicken breast."0.2s
Filter — controls the list below
Pass / Fail / Unverified
PASS 11FAIL 1UNVERIFIED 0
Module (intended for)
Food (12)
Utterance sub-type (within module)
12 shown — 11 pass, 1 fail, 0 unverified
Per-utterance detail
PASSshould log the entry"Log one medium banana for the sequence." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Banana — card not captured.
Why verdict Logged (write-truth verified): Banana — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Logged one medium banana for the sequence. Assumed 1 medium banana as the default serving. Tell me if you want a different amount.
3 · App data rows written created food_log_entry: Banana ×1 (1 medium banana) 105 cal · 1.3g P · 27.1g C · 0.4g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:27:25.139Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Track one Quest chocolate chip cookie dough bar." (difficulty —)0.5s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Quest chocolate chip cookie dough bar — card not captured.
Why verdict Logged (write-truth verified): Quest chocolate chip cookie dough bar — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 · TTS saidno speech captured
2 · Card shown Logged Quest chocolate chip cookie dough bar — the macros are an estimate since this food varies by brand.
3 · App data rows written created food_log_entry: Quest chocolate chip cookie dough bar ×1 (1 bar) 190 cal · 21g P · 22g C · 9g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:27:36.677Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from a manually curated nutrition entry.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
FAILshould log the entry"I had one cup cooked oatmeal with no toppings." (difficulty —)5.6s
Verdict Expected LOG — should log the entry. FAIL: OVER-ASK — asked instead of logging (no saved row).
Why verdict OVER-ASK — asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 5.6s
1 · TTS saidno speech captured
2 · Card shown What exact food and amount should I use for Cooked oatmeal? I did not log it yet because I could not safely finish that food log.
3 · App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:27:53.448Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Add two hard boiled eggs for lunch." (difficulty —)0.4s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Egg — card not captured.
Why verdict Logged (write-truth verified): Egg — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 · TTS saidno speech captured
2 · Card shown Logged two hard boiled eggs for lunch.
3 · App data rows written created food_log_entry: Egg ×1 (2 eggs) 143 cal · 12.6g P · 0.7g C · 9.5g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:05.131Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Log one cup blueberries as a snack." (difficulty —)0.3s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Blueberries — card not captured.
Why verdict Logged (write-truth verified): Blueberries — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.3s
1 · TTS saidno speech captured
2 · Card shown Logged one cup blueberries as a snack. Assumed 1 cup as the default serving. Tell me if you want a different amount.
3 · App data rows written created food_log_entry: Blueberries ×1 (1 cup) 84 cal · 1g P · 21.5g C · 0.4g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:16.496Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Record four ounces grilled chicken." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Chicken breast — card not captured.
Why verdict Logged (write-truth verified): Chicken breast — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Logged four ounces grilled chicken.
3 · App data rows written created food_log_entry: Chicken breast ×1 (4 oz) 187 cal · 35.2g P · 0g C · 4.1g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:27.948Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Track one tablespoon peanut butter." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Peanut butter — card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Logged one tablespoon peanut butter. Assumed 1 tbsp as the default serving. Tell me if you want a different amount.
3 · App data rows written created food_log_entry: Peanut butter ×1 (1 tbsp) 88 cal · 3.8g P · 2.9g C · 7.6g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:39.249Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I drank one cup two percent milk." (difficulty —)0.3s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): 2% milk — card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.3s
1 · TTS saidno speech captured
2 · Card shown Logged one cup two percent milk.
3 · App data rows written created food_log_entry: 2% milk ×1 (1 cup) 122 cal · 8.1g P · 11.7g C · 4.9g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:50.659Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had one medium banana." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Banana — card not captured.
Why verdict Logged (write-truth verified): Banana — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Logged one medium banana. Assumed 1 medium banana as the default serving. Tell me if you want a different amount.
3 · App data rows written created food_log_entry: Banana ×1 (1 medium banana) 105 cal · 1.3g P · 27.1g C · 0.4g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:29:01.929Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Breakfast was two hard boiled eggs." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Egg — card not captured.
Why verdict Logged (write-truth verified): Egg — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Logged two hard boiled eggs.
3 · App data rows written created food_log_entry: Egg ×1 (2 eggs) 143 cal · 12.6g P · 0.7g C · 9.5g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:29:13.281Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Put down four ounces grilled chicken breast." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Chicken breast — card not captured.
Why verdict Logged (write-truth verified): Chicken breast — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Logged four ounces grilled chicken breast.
3 · App data rows written created food_log_entry: Chicken breast ×1 (4 oz) 187 cal · 35.2g P · 0g C · 4.1g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:29:24.595Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Record one cup cooked white rice." (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Cooked white rice — card not captured.
Why verdict Logged (write-truth verified): Cooked white rice — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS saidno speech captured
2 · Card shown Logged one cup cooked white rice. Assumed 1 cup as the default serving. Tell me if you want a different amount.
3 · App data rows written created food_log_entry: Cooked white rice ×1 (1 cup) 205 cal · 4.3g P · 44.2g C · 0.5g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:29:35.983Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)