📊 Status Dashboard ↑ all runs

multifood-corpus-a-fb954a4d-c43-20260801

multifood-corpus-a · preserved partial run · 224s elapsed · iOS sim
This is the preserved partial result for this run. Its original utterance evidence is shown below.
Rows
14
Pass
8 (57%)
Fail
5 (36%)
Unverified
1 (7%)
Pass rate
62%
Avg difficulty
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace — needs a human look (that's you 👍/👎-ing it). "Pass rate" = pass ÷ (pass + fail) — it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows). Unverified breaks down as: 1 unclassified — the "grader coverage gap" ones aren't ambiguous, the grader just doesn't fully score that intent type yet.

Why the fails happened — comprehension vs execution vs cosmetic

Comprehension — picked the wrong action/target (the hard problem)
5 (100%)
Of 5 fails: if most are comprehension, that's the hard problem (the app didn't figure out the right action/target); execution/cosmetic means the decision was right but something downstream broke. Click a bar for the sub-split + example utterances.

Handled correctly? — by expected action

Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed toNCorrectWrongUnverified
▸ LOG — log the entry 104 (40%) 5 (50%) 1 (10%)
▸ CLARIFY — ask a clarifying question 44 (100%) 0 (0%) 0 (0%)
Total148 (62%)51

Accuracy by difficulty

Pending A1's per-utterance difficulty score (requested 2026-07-05) — this bar chart lights up once that lands.

Clarification follow-ups — scored separately

Second turn: app asked, we replied — did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.

Cosmetic only

Not yet classified — pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.

System / infra

Not yet classified — pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.

Latency

Avg (time to ready)
1.6s
p90
3.2s
Max
3.7s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers — click any dot to jump to its detail.
0s
1s
2s
5s
4s
Response path — quick (single response) vs async (an ack like "Working on it…" before the real answer).
Quick response
10
Sync clarification
4
Slowest 8 utterances (click to jump to detail):
"Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice."3.7s
"I had a Chipotle chicken burrito bowl and a bottle of Topo Chico."3.2s
"I ate a turkey sandwich on rye, some chips, and a Diet Coke."3.2s
"Breakfast was two scrambled eggs, one slice of sourdough toast, and eight ounces of cold brew."3.2s
"Snack spread: one Blue Diamond smokehouse almonds pack, one Babybel light, one cup edamame, one Pink Lady apple, one cup coconut water, one Good Culture strawberry cottage cheese cup, one rice cake, one tablespoon chia seeds, five walnuts, and one Clif Kid Zbar."2.1s
"Log a Starbucks tall oat milk latte, one Kodiak Cakes blueberry waffle, two strips turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and one Justin maple almond butter squeeze pack."1.4s
"Dinner: eight ounces lean ground turkey, one cup roasted Brussels sprouts, half avocado, one Thomas whole wheat English muffin, one Laughing Cow wedge, ten baby carrots, and sixteen ounces sparkling water."1.3s
"I ate a Panera Mediterranean bowl, one Perfect Bar almond butter, a cup of bone broth, two celery stalks, and one Medjool date."1.1s

Filter — controls the list below

Pass / Fail / Unverified
PASS 8 FAIL 5 UNVERIFIED 1
Module (intended for)
food (14)
Utterance sub-type (within module)
14 shown — 8 pass, 5 fail, 1 unverified

Per-utterance detail

FAILshould log the entry"Log a medium Fuji apple and one KIND peanut butter dark chocolate bar." (difficulty —)0.9s
Verdict Expected LOG — should log the entry. FAIL: WRITE-TRUTH FAIL — WRONG/MISSING item "KIND peanut butter dark chocolate bar" — no saved row with matching identity (rows: Chocolate Candy, Peanut Butter Filled, Apple)
Why verdict WRITE-TRUTH FAIL — WRONG/MISSING item "KIND peanut butter dark chocolate bar" — no saved row with matching identity (rows: Chocolate Candy, Peanut Butter Filled, Apple)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.9s
1 · TTS said no speech captured
2 · Card shown Logged Apple and Chocolate Candy, Peanut Butter Filled. Assumed catalog default servings where you did not say an amount — tell me if that is not right.
3 · App data rows written created food_log_entry: Apple ×1 (182 g) 95 cal · 0.5g P · 25.5g C · 0.4g F
created food_log_entry: Chocolate Candy, Peanut Butter Filled ×1 (1 large/king size) 438 cal · 7g P · 57.7g C · 22.2g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:26:50.805Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
FAILshould log the entry"I had a Chipotle chicken burrito bowl and a bottle of Topo Chico." (difficulty —)3.2s
Verdict Expected LOG — should log the entry. FAIL: OVER-ASK — asked instead of logging (no saved row).
Why verdict OVER-ASK — asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.2s
1 · TTS said no speech captured
2 · Card shown I need to resolve a Chipotle chicken burrito bowl before I log this meal. What should I use for a Chipotle chicken burrito bowl?
3 · App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:27:05.171Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method No lookup method was captured for this path.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
FAILshould log the entry"Breakfast was two scrambled eggs, one slice of sourdough toast, and eight ounces of cold brew." (difficulty —)3.2s
Verdict Expected LOG — should log the entry. FAIL: OVER-ASK — asked instead of logging (no saved row).
Why verdict OVER-ASK — asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.2s
1 · TTS said no speech captured
2 · Card shown I need to resolve eight ounces of cold brew before I log this meal. What should I use for eight ounces of cold brew?
3 · App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:27:19.547Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method No lookup method was captured for this path.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
FAILshould log the entry"Track one Oikos triple zero vanilla yogurt, a medium navel orange, and twelve almonds." (difficulty —)0.7s
Verdict Expected LOG — should log the entry. FAIL: WRITE-TRUTH FAIL — WRONG/MISSING item "Oikos triple zero vanilla yogurt" — no saved row with matching identity (rows: Plain Greek yogurt, Orange, Almonds)
Why verdict WRITE-TRUTH FAIL — WRONG/MISSING item "Oikos triple zero vanilla yogurt" — no saved row with matching identity (rows: Plain Greek yogurt, Orange, Almonds)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.7s
1 · TTS said no speech captured
2 · Card shown Logged Plain Greek yogurt, Orange, and twelve almonds. Assumed catalog default servings where you did not say an amount — tell me if that is not right.
3 · App data rows written created food_log_entry: Plain Greek yogurt ×1 (100 g) 97 cal · 9g P · 3.9g C · 5g F
created food_log_entry: Orange ×1 (184 g) 86 cal · 1.7g P · 21.7g C · 0.2g F
created food_log_entry: Almonds ×1 (12 almonds) 83 cal · 3.1g P · 3.1g C · 7.2g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:27:31.566Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Lunch was six ounces grilled Atlantic salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge." (difficulty —)0.6s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Salmon, Broccoli, Cooked brown rice, Olive oil, Lemon — card not captured.
Why verdict Logged (write-truth verified): Salmon, Broccoli, Cooked brown rice, Olive oil, Lemon — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 · TTS said no speech captured
2 · Card shown Logged Salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 · App data rows written created food_log_entry: Salmon ×1 (six ounces (170.1 g)) 354 cal · 34g P · 0g C · 22.1g F
created food_log_entry: Broccoli ×1 (1 cup) 55 cal · 3.8g P · 11.3g C · 0.6g F
created food_log_entry: Cooked brown rice ×1 (0.5 cup) 109 cal · 2.2g P · 22.9g C · 0.8g F
created food_log_entry: Olive oil ×1 (1 tbsp) 120 cal · 0g P · 0g C · 13.6g F
created food_log_entry: Lemon ×1 (1 lemon wedge) 17 cal · 0.6g P · 5.4g C · 0.2g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:27:43.325Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
FAILshould log the entry"I ate a Panera Mediterranean bowl, one Perfect Bar almond butter, a cup of bone broth, two celery stalks, and one Medjool date." (difficulty —)1.1s
Verdict Expected LOG — should log the entry. FAIL: WRITE-TRUTH FAIL — IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving])
Why verdict WRITE-TRUTH FAIL — IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving])
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.1s
1 · TTS said no speech captured
2 · Card shown Logged Panera Bread Mediterranean Bowl, Almond Butter — PERFECT BAR, a cup of bone broth, two celery stalks, and one medjool date. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 · App data rows written created food_log_entry: Panera Bread Mediterranean Bowl ×1 (1 Bowl) 590 cal · 16g P · 66g C · 31g F
created food_log_entry: Almond Butter ×1 (1 BAR) 320 cal · 13g P · 25g C · 19g F
created food_log_entry: Bone broth ×1 (1 cup) 36 cal · 7.2g P · 1.2g C · 0.7g F
created food_log_entry: Celery ×1 (2 celery stalks) 13 cal · 0.6g P · 2.4g C · 0.2g F
created food_log_entry: Dates ×1 (1 medjool date) 66 cal · 0.4g P · 18g C · 0g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:27:55.750Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from restaurant menu.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Dinner: eight ounces lean ground turkey, one cup roasted Brussels sprouts, half avocado, one Thomas whole wheat English muffin, one Laughing Cow wedge, ten baby carrots, and sixteen ounces sparkling water." (difficulty —)1.3s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Baby carrots, Water, Brussels sprouts, Lean ground turkey, Avocado, Whole wheat English muffin, Laughing Cow — card not captured.
Why verdict Logged (write-truth verified): Baby carrots, Water, Brussels sprouts, Lean ground turkey, Avocado, Whole wheat English muffin, Laughing Cow — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 · TTS said no speech captured
2 · Card shown Logged eight ounces lean ground turkey, one cup roasted brussels sprouts, half avocado, Whole wheat English muffin, one laughing cow wedge, ten baby carrots, and Water. Assumed catalog default servings where you did not say an amount — tell me if that is not right.
3 · App data rows written created food_log_entry: Lean ground turkey ×1 (8 oz) 386 cal · 61.2g P · 0g C · 15.9g F
created food_log_entry: Brussels sprouts ×1 (1 cup) 70 cal · 5.4g P · 14.4g C · 0.5g F
created food_log_entry: Avocado ×1 (half avocado) 120 cal · 1.5g P · 6.4g C · 11g F
created food_log_entry: Whole wheat English muffin ×1 (1 muffin (57 g)) 128 cal · 5.6g P · 24.8g C · 1.3g F
created food_log_entry: Laughing Cow ×1 (1 laughing cow wedge) 35 cal · 2g P · 1g C · 2.5g F
created food_log_entry: Baby carrots ×1 (10 baby carrots) 35 cal · 0.6g P · 8.2g C · 0.1g F
created food_log_entry: Water ×1 (sixteen ounces (453.6 g)) 0 cal · 0g P · 0g C · 0g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:08.261Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Log a Starbucks tall oat milk latte, one Kodiak Cakes blueberry waffle, two strips turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and one Justin maple almond butter squeeze pack." (difficulty —)1.4s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Oat milk latte, Turkey bacon, Blackberries, Almond butter, Blueberry waffle, Egg, Mozzarella — card not captured.
Why verdict Logged (write-truth verified): Oat milk latte, Turkey bacon, Blackberries, Almond butter, Blueberry waffle, Egg, Mozzarella — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.4s
1 · TTS said no speech captured
2 · Card shown Logged a starbucks tall oat milk latte, one kodiak cakes blueberry waffle, Turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and Almond butter. Assumed catalog default servings where you did not say an amount — tell me if that is not right.
3 · App data rows written created food_log_entry: Oat milk latte ×1 (1 starbucks tall oat milk latte) 135 cal · 4.3g P · 18.5g C · 5.3g F
created food_log_entry: Blueberry waffle ×1 (1 kodiak cakes blueberry waffle) 180 cal · 7g P · 24.5g C · 4.9g F
created food_log_entry: Turkey bacon ×2 (14 g) 60 cal · 8.2g P · 1g C · 3g F
created food_log_entry: Blackberries ×1 (1 cup) 62 cal · 2g P · 14.7g C · 0.7g F
created food_log_entry: Egg ×1 (1 egg) 72 cal · 6.3g P · 0.4g C · 4.8g F
created food_log_entry: Mozzarella ×1 (1 oz) 72 cal · 6.9g P · 0.8g C · 4.5g F
created food_log_entry: Almond butter ×1 (32 g) 196 cal · 6.7g P · 6.1g C · 17.9g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:20.858Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Snack spread: one Blue Diamond smokehouse almonds pack, one Babybel light, one cup edamame, one Pink Lady apple, one cup coconut water, one Good Culture strawberry cottage cheese cup, one rice cake, one tablespoon chia seeds, five walnuts, and one Clif Kid Zbar." (difficulty —)2.1s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Almonds, Chia seeds, Plain rice cakes, Walnuts, Babybel light, Apple, Cottage cheese, Coconut water, Edamame, Clif Kid Zbar — card not captured.
Why verdict Logged (write-truth verified): Almonds, Chia seeds, Plain rice cakes, Walnuts, Babybel light, Apple, Cottage cheese, Coconut water, Edamame, Clif Kid Zbar — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.1s
1 · TTS said no speech captured
2 · Card shown Logged Almonds, one babybel light, one cup edamame, Apple, one cup coconut water, Cottage cheese, one rice cake, one tablespoon chia seeds, five walnuts, and one clif kid zbar. Assumed catalog default servings where you did not say an amount — tell me if that is not right.
3 · App data rows written created food_log_entry: Almonds ×1 (28 g) 162 cal · 5.9g P · 6g C · 14g F
created food_log_entry: Babybel light ×1 (1 babybel light) 42 cal · 6g P · 0g C · 2.5g F
created food_log_entry: Edamame ×1 (1 cup) 188 cal · 18.4g P · 13.8g C · 8.1g F
created food_log_entry: Apple ×1 (182 g) 95 cal · 0.5g P · 25.5g C · 0.4g F
created food_log_entry: Coconut water ×1 (1 cup) 46 cal · 1.7g P · 8.9g C · 0.5g F
created food_log_entry: Cottage cheese ×1 (1 cup) 221 cal · 25.1g P · 7.7g C · 9.7g F
created food_log_entry: Plain rice cakes ×1 (1 rice cake) 35 cal · 0.7g P · 7.3g C · 0.3g F
created food_log_entry: Chia seeds ×1 (1 tbsp) 51 cal · 1.7g P · 4.4g C · 3.2g F
created food_log_entry: Walnuts ×1 (5 walnuts) 65 cal · 1.5g P · 1.4g C · 6.5g F
created food_log_entry: Clif Kid Zbar ×1 (1 clif kid zbar) 130 cal · 3g P · 23g C · 4g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:34.035Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice." (difficulty —)3.7s
Verdict Expected LOG — should log the entry. UNVERIFIED: UNVERIFIABLE — no saved row captured (untagged / lossy) — resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE — no saved row captured (untagged / lossy) — resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 · TTS said no speech captured
2 · Card shown Logging your Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice — still looking up the nutrition details; they'll appear in your log shortly.
3 · App data rows written pending_created food_log_entry: Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice (pending)
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:28:55.606Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method No lookup method was captured for this path.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould ask a clarifying question"I had toast and a Chobani peach Greek yogurt." (difficulty —)0.4s
Verdict Expected CLARIFY — should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 · TTS said no speech captured
2 · Card shown Logged Toast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount — tell me if that is not right.
3 · App data rows written created food_log_entry: Toast ×1 (100 g) 265 cal · 9g P · 49g C · 3.2g F
created food_log_entry: Plain Greek yogurt ×1 (100 g) 97 cal · 9g P · 3.9g C · 5g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:29:00.445Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould ask a clarifying question"Log cottage cheese and one large Honeycrisp apple." (difficulty —)0.6s
Verdict Expected CLARIFY — should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 · TTS said no speech captured
2 · Card shown Logged Cottage cheese and Apple. Assumed catalog default servings where you did not say an amount — tell me if that is not right.
3 · App data rows written created food_log_entry: Cottage cheese ×1 (100 g) 98 cal · 11.1g P · 3.4g C · 4.3g F
created food_log_entry: Apple ×1 (182 g) 95 cal · 0.5g P · 25.5g C · 0.4g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:29:12.278Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from the common-food list.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould ask a clarifying question"I ate a turkey sandwich on rye, some chips, and a Diet Coke." (difficulty —)3.2s
Verdict Expected CLARIFY — should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.2s
1 · TTS said no speech captured
2 · Card shown I need to resolve a turkey sandwich on rye before I log this meal. What should I use for a turkey sandwich on rye?
3 · App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:29:26.621Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method No lookup method was captured for this path.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould ask a clarifying question"Breakfast was two eggs, toast, and black coffee." (difficulty —)0.1s
Verdict Expected CLARIFY — should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.1s
1 · TTS said no speech captured
2 · Card shown I did not log that yet because the exact item and amount decide the nutrition. What food and how much should I log?
3 · App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:29:37.960Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method No lookup method was captured for this path.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)