πŸ“Š Status Dashboard ↑ all runs

multifood-corpus-a-14385762-c55-20260802

multifood-corpus-a Β· 2 minutes ago Β· iOS sim
Rows
60
Pass
36 (60%)
Fail
17 (28%)
Unverified
7 (12%)
Pass rate
68%
Avg difficulty
β€”
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace β€” needs a human look (that's you πŸ‘/πŸ‘Ž-ing it). "Pass rate" = pass Γ· (pass + fail) β€” it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows). Unverified breaks down as: 7 unclassified β€” the "grader coverage gap" ones aren't ambiguous, the grader just doesn't fully score that intent type yet.

Why the fails happened β€” comprehension vs execution vs cosmetic

Comprehension β€” picked the wrong action/target (the hard problem)
17 (100%)
Of 17 fails: if most are comprehension, that's the hard problem (the app didn't figure out the right action/target); execution/cosmetic means the decision was right but something downstream broke. Click a bar for the sub-split + example utterances.

Handled correctly? β€” by expected action

Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed toNCorrectWrongUnverified
β–Έ CLARIFY β€” ask a clarifying question 3933 (85%) 6 (15%) 0 (0%)
β–Έ LOG β€” log the entry 213 (14%) 11 (52%) 7 (33%)
Total6036 (68%)177

Accuracy by difficulty

Pending A1's per-utterance difficulty score (requested 2026-07-05) β€” this bar chart lights up once that lands.

Clarification follow-ups β€” scored separately

Second turn: app asked, we replied β€” did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.

Cosmetic only

Not yet classified β€” pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.

System / infra

Not yet classified β€” pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.

Latency

Avg (time to ready)
2.1s
p90
3.7s
Max
5.5s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers β€” click any dot to jump to its detail.
0s
1s
2s
5s
6s
Response path β€” quick (single response) vs async (an ack like "Working on it…" before the real answer).
Sync clarification
39
Quick response
21
Slowest 8 utterances (click to jump to detail):
"Track one cup overnight oats with chia, one scoop Vital Proteins collagen peptides, one cup soy milk, one tablespoon maple syrup, one medium blood orange, one ounce hemp hearts, and one cup plain kefir."5.5s
"Meal: one Impossible Whopper from Burger King, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard."5.3s
"Lunch was a Chipotle steak burrito, granola, one Fuji apple, one cup sparkling water, and a string cheese."4.5s
"I ate one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream."3.9s
"Log one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar."3.7s
"Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana."3.7s
"Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig."3.7s
"I ate pizza, one cup Greek salad with feta, one bottle La Croix lime, two eggs, one cup vanilla bean ice cream, one slice whole grain bread, and six ounces Atlantic cod."3.7s

Filter β€” controls the list below

Pass / Fail / Unverified
PASS 36 FAIL 17 UNVERIFIED 7
Module (intended for)
food (60)
Utterance sub-type (within module)
60 shown β€” 36 pass, 17 fail, 7 unverified

Per-utterance detail

FAILshould log the entry"Log a medium Fuji apple and one KIND peanut butter dark chocolate bar." (difficulty β€”)0.8s
Verdict Expected LOG β€” should log the entry. FAIL: WRITE-TRUTH FAIL β€” WRONG/MISSING item "KIND peanut butter dark chocolate bar" β€” no saved row with matching identity (rows: Apple, Chocolate Candy, Peanut Butter Filled)
Why verdict WRITE-TRUTH FAIL β€” WRONG/MISSING item "KIND peanut butter dark chocolate bar" β€” no saved row with matching identity (rows: Apple, Chocolate Candy, Peanut Butter Filled)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.8s
1 Β· TTS said Logged Apple and Chocolate Candy, Peanut Butter Filled. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Apple and Chocolate Candy, Peanut Butter Filled. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
created food_log_entry: Chocolate Candy, Peanut Butter Filled Γ—1 (1 large/king size) 438 cal Β· 7g P Β· 57.7g C Β· 22.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:47:06.673Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"I had a Chipotle chicken burrito bowl and a bottle of Topo Chico." (difficulty β€”)3.1s
Verdict Expected LOG β€” should log the entry. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.1s
1 Β· TTS said I need to resolve a Chipotle chicken burrito bowl before I log this meal. What should I use for a Chipotle chicken burrito bowl?
2 Β· Card shown I need to resolve a Chipotle chicken burrito bowl before I log this meal. What should I use for a Chipotle chicken burrito bowl?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:47:20.900Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"Breakfast was two scrambled eggs, one slice of sourdough toast, and eight ounces of cold brew." (difficulty β€”)2.9s
Verdict Expected LOG β€” should log the entry. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.9s
1 Β· TTS said I need to resolve eight ounces of cold brew before I log this meal. What should I use for eight ounces of cold brew?
2 Β· Card shown I need to resolve eight ounces of cold brew before I log this meal. What should I use for eight ounces of cold brew?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:47:34.978Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"Track one Oikos triple zero vanilla yogurt, a medium navel orange, and twelve almonds." (difficulty β€”)0.7s
Verdict Expected LOG β€” should log the entry. FAIL: WRITE-TRUTH FAIL β€” WRONG/MISSING item "Oikos triple zero vanilla yogurt" β€” no saved row with matching identity (rows: Plain Greek yogurt, Orange, Almonds)
Why verdict WRITE-TRUTH FAIL β€” WRONG/MISSING item "Oikos triple zero vanilla yogurt" β€” no saved row with matching identity (rows: Plain Greek yogurt, Orange, Almonds)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.7s
1 Β· TTS said Logged Plain Greek yogurt, Orange, and twelve almonds. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Plain Greek yogurt, Orange, and twelve almonds. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Plain Greek yogurt Γ—1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
created food_log_entry: Orange Γ—1 (184 g) 86 cal Β· 1.7g P Β· 21.7g C Β· 0.2g F
created food_log_entry: Almonds Γ—1 (12 almonds) 83 cal Β· 3.1g P Β· 3.1g C Β· 7.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:47:46.849Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the entry"Lunch was six ounces grilled Atlantic salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge." (difficulty β€”)0.6s
Verdict Expected LOG β€” should log the entry. PASS: Logged (write-truth verified): Broccoli, Salmon, Lemon, Olive oil, Cooked brown rice β€” card not captured.
Why verdict Logged (write-truth verified): Broccoli, Salmon, Lemon, Olive oil, Cooked brown rice β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS said Logged Salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
2 Β· Card shown Logged Salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Salmon Γ—1 (six ounces (170.1 g)) 354 cal Β· 34g P Β· 0g C Β· 22.1g F
created food_log_entry: Broccoli Γ—1 (1 cup) 55 cal Β· 3.8g P Β· 11.3g C Β· 0.6g F
created food_log_entry: Cooked brown rice Γ—1 (0.5 cup) 109 cal Β· 2.2g P Β· 22.9g C Β· 0.8g F
created food_log_entry: Olive oil Γ—1 (1 tbsp) 120 cal Β· 0g P Β· 0g C Β· 13.6g F
created food_log_entry: Lemon Γ—1 (1 lemon wedge) 17 cal Β· 0.6g P Β· 5.4g C Β· 0.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:47:58.562Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"I ate a Panera Mediterranean bowl, one Perfect Bar almond butter, a cup of bone broth, two celery stalks, and one Medjool date." (difficulty β€”)1.0s
Verdict Expected LOG β€” should log the entry. FAIL: WRITE-TRUTH FAIL β€” IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving]); WRONG/MISSING item "Medjool date" β€” no saved row with matching identity (rows: Dates, Panera Bread Mediterranean Bowl, Almond Butter, Almond Butter, Celery, Bone broth, Celery, Dates)
Why verdict WRITE-TRUTH FAIL β€” IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving]); WRONG/MISSING item "Medjool date" β€” no saved row with matching identity (rows: Dates, Panera Bread Mediterranean Bowl, Almond Butter, Almond Butter, Celery, Bone broth, Celery, Dates)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS said Logged Panera Bread Mediterranean Bowl, Almond Butter β€” PERFECT BAR, a cup of bone broth, two celery stalks, and one medjool date. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
2 Β· Card shown Logged Panera Bread Mediterranean Bowl, Almond Butter β€” PERFECT BAR, a cup of bone broth, two celery stalks, and one medjool date. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Panera Bread Mediterranean Bowl Γ—1 (1 Bowl) 590 cal Β· 16g P Β· 66g C Β· 31g F
created food_log_entry: Almond Butter Γ—1 (1 BAR) 320 cal Β· 13g P Β· 25g C Β· 19g F
created food_log_entry: Bone broth Γ—1 (1 cup) 36 cal Β· 7.2g P Β· 1.2g C Β· 0.7g F
created food_log_entry: Celery Γ—1 (2 celery stalks) 13 cal Β· 0.6g P Β· 2.4g C Β· 0.2g F
created food_log_entry: Dates Γ—1 (1 medjool date) 66 cal Β· 0.4g P Β· 18g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:48:10.699Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the entry"Dinner: eight ounces lean ground turkey, one cup roasted Brussels sprouts, half avocado, one Thomas whole wheat English muffin, one Laughing Cow wedge, ten baby carrots, and sixteen ounces sparkling water." (difficulty β€”)1.3s
Verdict Expected LOG β€” should log the entry. PASS: Logged (write-truth verified): Water, Whole wheat English muffin, Avocado, Brussels sprouts, Baby carrots, Lean ground turkey, Water, Laughing Cow β€” card not captured.
Why verdict Logged (write-truth verified): Water, Whole wheat English muffin, Avocado, Brussels sprouts, Baby carrots, Lean ground turkey, Water, Laughing Cow β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 Β· TTS said Logged eight ounces lean ground turkey, one cup roasted brussels sprouts, half avocado, Whole wheat English muffin, one laughing cow wedge, ten baby carrots, and Water. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged eight ounces lean ground turkey, one cup roasted brussels sprouts, half avocado, Whole wheat English muffin, one laughing cow wedge, ten baby carrots, and Water. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Lean ground turkey Γ—1 (8 oz) 386 cal Β· 61.2g P Β· 0g C Β· 15.9g F
created food_log_entry: Brussels sprouts Γ—1 (1 cup) 70 cal Β· 5.4g P Β· 14.4g C Β· 0.5g F
created food_log_entry: Avocado Γ—1 (half avocado) 120 cal Β· 1.5g P Β· 6.4g C Β· 11g F
created food_log_entry: Whole wheat English muffin Γ—1 (1 muffin (57 g)) 128 cal Β· 5.6g P Β· 24.8g C Β· 1.3g F
created food_log_entry: Laughing Cow Γ—1 (1 laughing cow wedge) 35 cal Β· 2g P Β· 1g C Β· 2.5g F
created food_log_entry: Baby carrots Γ—1 (10 baby carrots) 35 cal Β· 0.6g P Β· 8.2g C Β· 0.1g F
created food_log_entry: Water Γ—1 (sixteen ounces (453.6 g)) 0 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:48:23.173Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"Log a Starbucks tall oat milk latte, one Kodiak Cakes blueberry waffle, two strips turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and one Justin maple almond butter squeeze pack." (difficulty β€”)1.3s
Verdict Expected LOG β€” should log the entry. FAIL: WRITE-TRUTH FAIL β€” WRONG/MISSING item "one ounce part skim mozzarella" β€” no saved row with matching identity (rows: Blackberries, Egg, Turkey bacon, Blueberry waffle, Oat milk latte); WRONG/MISSING item "Justin maple almond butter squeeze pack" β€” no saved row with matching identity (rows: Blackberries, Egg, Turkey bacon, Blueberry waffle, Oat milk latte)
Why verdict WRITE-TRUTH FAIL β€” WRONG/MISSING item "one ounce part skim mozzarella" β€” no saved row with matching identity (rows: Blackberries, Egg, Turkey bacon, Blueberry waffle, Oat milk latte); WRONG/MISSING item "Justin maple almond butter squeeze pack" β€” no saved row with matching identity (rows: Blackberries, Egg, Turkey bacon, Blueberry waffle, Oat milk latte)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 Β· TTS said Logged a starbucks tall oat milk latte, one kodiak cakes blueberry waffle, Turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and Almond butter. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged a starbucks tall oat milk latte, one kodiak cakes blueberry waffle, Turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and Almond butter. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Oat milk latte Γ—1 (1 starbucks tall oat milk latte) 135 cal Β· 4.3g P Β· 18.5g C Β· 5.3g F
created food_log_entry: Blueberry waffle Γ—1 (1 kodiak cakes blueberry waffle) 180 cal Β· 7g P Β· 24.5g C Β· 4.9g F
created food_log_entry: Turkey bacon Γ—2 (14 g) 60 cal Β· 8.2g P Β· 1g C Β· 3g F
created food_log_entry: Blackberries Γ—1 (1 cup) 62 cal Β· 2g P Β· 14.7g C Β· 0.7g F
created food_log_entry: Egg Γ—1 (1 egg) 72 cal Β· 6.3g P Β· 0.4g C Β· 4.8g F
created food_log_entry: Mozzarella Γ—1 (1 oz) 72 cal Β· 6.9g P Β· 0.8g C Β· 4.5g F
created food_log_entry: Almond butter Γ—1 (32 g) 196 cal Β· 6.7g P Β· 6.1g C Β· 17.9g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:48:35.692Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"Snack spread: one Blue Diamond smokehouse almonds pack, one Babybel light, one cup edamame, one Pink Lady apple, one cup coconut water, one Good Culture strawberry cottage cheese cup, one rice cake, one tablespoon chia seeds, five walnuts, and one Clif Kid Zbar." (difficulty β€”)2.1s
Verdict Expected LOG β€” should log the entry. FAIL: WRITE-TRUTH FAIL β€” WRONG/MISSING item "one cup coconut water" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "Good Culture strawberry cottage cheese cup" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "rice cake" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "one tablespoon chia seeds" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "five walnuts" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "Clif Kid Zbar" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds)
Why verdict WRITE-TRUTH FAIL β€” WRONG/MISSING item "one cup coconut water" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "Good Culture strawberry cottage cheese cup" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "rice cake" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "one tablespoon chia seeds" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "five walnuts" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds); WRONG/MISSING item "Clif Kid Zbar" β€” no saved row with matching identity (rows: Edamame, Babybel light, Apple, Almonds)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.1s
1 Β· TTS said Logged Almonds, one babybel light, one cup edamame, Apple, one cup coconut water, Cottage cheese, one rice cake, one tablespoon chia seeds, five walnuts, and one clif kid zbar. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Almonds, one babybel light, one cup edamame, Apple, one cup coconut water, Cottage cheese, one rice cake, one tablespoon chia seeds, five walnuts, and one clif kid zbar. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Almonds Γ—1 (28 g) 162 cal Β· 5.9g P Β· 6g C Β· 14g F
created food_log_entry: Babybel light Γ—1 (1 babybel light) 42 cal Β· 6g P Β· 0g C Β· 2.5g F
created food_log_entry: Edamame Γ—1 (1 cup) 188 cal Β· 18.4g P Β· 13.8g C Β· 8.1g F
created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
created food_log_entry: Coconut water Γ—1 (1 cup) 46 cal Β· 1.7g P Β· 8.9g C Β· 0.5g F
created food_log_entry: Cottage cheese Γ—1 (1 cup) 221 cal Β· 25.1g P Β· 7.7g C Β· 9.7g F
created food_log_entry: Plain rice cakes Γ—1 (1 rice cake) 35 cal Β· 0.7g P Β· 7.3g C Β· 0.3g F
created food_log_entry: Chia seeds Γ—1 (1 tbsp) 51 cal Β· 1.7g P Β· 4.4g C Β· 3.2g F
created food_log_entry: Walnuts Γ—1 (5 walnuts) 65 cal Β· 1.5g P Β· 1.4g C Β· 6.5g F
created food_log_entry: Clif Kid Zbar Γ—1 (1 clif kid zbar) 130 cal Β· 3g P Β· 23g C Β· 4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:48:48.899Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice." (difficulty β€”)3.7s
Verdict Expected LOG β€” should log the entry. UNVERIFIED: UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS said Logging your Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice β€” still looking up the nutrition details; they'll appear in your log shortly.
2 Β· Card shown Logging your Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice β€” still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice (pending)
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:49:11.099Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had toast and a Chobani peach Greek yogurt." (difficulty β€”)0.5s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 Β· TTS said Logged Toast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Toast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Toast Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
created food_log_entry: Plain Greek yogurt Γ—1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:49:15.293Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Log cottage cheese and one large Honeycrisp apple." (difficulty β€”)0.6s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS said Logged Cottage cheese and Apple. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Cottage cheese and Apple. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Cottage cheese Γ—1 (100 g) 98 cal Β· 11.1g P Β· 3.4g C Β· 4.3g F
created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:49:27.034Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I ate a turkey sandwich on rye, some chips, and a Diet Coke." (difficulty β€”)3.1s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.1s
1 Β· TTS said I need to resolve a turkey sandwich on rye before I log this meal. What should I use for a turkey sandwich on rye?
2 Β· Card shown I need to resolve a turkey sandwich on rye before I log this meal. What should I use for a turkey sandwich on rye?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:49:41.312Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Breakfast was two eggs, toast, and black coffee." (difficulty β€”)0.1s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.1s
1 Β· TTS said I did not log that yet because the exact item and amount decide the nutrition. What food and how much should I log?
2 Β· Card shown I did not log that yet because the exact item and amount decide the nutrition. What food and how much should I log?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:49:52.489Z
6 Β· Why this food
  • The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • Matched your saved foods and aliases to "toast" at 90% confidence.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 0.8 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had a protein shake, one RXBAR chocolate sea salt bar, and a banana." (difficulty β€”)0.8s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.8s
1 Β· TTS said I need to resolve one RXBAR chocolate sea salt bar before I log this meal. What should I use for one RXBAR chocolate sea salt bar?
2 Β· Card shown I need to resolve one RXBAR chocolate sea salt bar before I log this meal. What should I use for one RXBAR chocolate sea salt bar?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:50:04.441Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Lunch was a Chipotle steak burrito, granola, one Fuji apple, one cup sparkling water, and a string cheese." (difficulty β€”)4.5s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 4.5s
1 Β· TTS said I need to resolve a Chipotle steak burrito before I log this meal. What should I use for a Chipotle steak burrito?
2 Β· Card shown I need to resolve a Chipotle steak burrito before I log this meal. What should I use for a Chipotle steak burrito?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:50:20.123Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had six ounces grilled salmon, peanut butter, one cup steamed asparagus, one Diet Pepsi, and one ounce sharp cheddar." (difficulty β€”)0.7s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.7s
1 Β· TTS said Logged six ounces grilled salmon, Peanut butter, one cup steamed asparagus, Diet cola, and Cheddar cheese. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged six ounces grilled salmon, Peanut butter, one cup steamed asparagus, Diet cola, and Cheddar cheese. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Salmon Γ—1 (6 oz) 354 cal Β· 34g P Β· 0g C Β· 22.1g F
created food_log_entry: Peanut butter Γ—1 (32 g) 188 cal Β· 8g P Β· 6.3g C Β· 16.1g F
created food_log_entry: Asparagus Γ—1 (1 cup) 40 cal Β· 4.3g P Β· 7.4g C Β· 0.4g F
created food_log_entry: Diet cola Γ—1 (355 g) 0 cal Β· 0g P Β· 0g C Β· 0g F
created food_log_entry: Cheddar cheese Γ—1 (one ounce (28.3 g)) 114 cal Β· 6.5g P Β· 1g C Β· 9.4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:50:31.914Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Track a Wendy's spicy chicken sandwich, cereal, one cup two percent milk, one clementine, and one ounce provolone." (difficulty β€”)1.8s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.8s
1 Β· TTS said I need to resolve cereal before I log this meal. What should I use for cereal?
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:50:44.872Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould ask a clarifying question"Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent." (difficulty β€”)3.6s
Verdict Expected CLARIFY β€” should ask a clarifying question. FAIL: Logged a BLIND guess β€” no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β€” no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS said Logging your Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent β€” still looking up the nutrition details; they'll appear in your log shortly.
2 Β· Card shown Logging your Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent β€” still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent (pending)
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:51:06.153Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I ate pizza, one cup Greek salad with feta, one bottle La Croix lime, two eggs, one cup vanilla bean ice cream, one slice whole grain bread, and six ounces Atlantic cod." (difficulty β€”)3.7s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS said I need to resolve one cup vanilla bean ice cream before I log this meal. What should I use for one cup vanilla bean ice cream?
2 Β· Card shown I need to resolve one cup vanilla bean ice cream before I log this meal. What should I use for one cup vanilla bean ice cream?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:51:14.422Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Log yogurt, one medium Bartlett pear, one ounce pistachios, one In-N-Out grilled cheese, one tablespoon Jif creamy peanut butter, one cup peppermint tea, and one mandarin." (difficulty β€”)3.3s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.3s
1 Β· TTS said I need to resolve one tablespoon Jif creamy peanut butter, one cup peppermint tea, and one mandarin before I log this meal. What should I use for one tablespoon Jif creamy peanut butter, one cup peppermint tea, and one mandarin?
2 Β· Card shown I need to resolve one tablespoon Jif creamy peanut butter, one cup peppermint tea, and one mandarin before I log this meal. What should I use for one tablespoon Jif creamy peanut butter, one cup peppermint tea, and one mandarin?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:51:40.346Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Big breakfast: two eggs, one strip turkey bacon, one medium banana, one cup black coffee, one slice whole wheat bread, six ounces plain nonfat Greek yogurt, and toast." (difficulty β€”)1.1s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.1s
1 Β· TTS said Logged Turkey bacon. Using your recent Turkey bacon history. Tell me if that is wrong.
2 Β· Card shown Logged Turkey bacon. Using your recent Turkey bacon history. Tell me if that is wrong.
3 Β· App data rows written created food_log_entry: Turkey bacon Γ—0.14285714285714285 (14 g) 4 cal Β· 0.6g P Β· 0.1g C Β· 0.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:51:52.618Z
6 Β· Why this food
  • The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • Matched your saved foods and aliases to "turkey bacon" at 90% confidence.
  • Matched your past food logs to "Toast" at 82% confidence.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 2.4 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Meal prep: four ounces chicken breast, one cup jasmine rice, one cup green beans, hummus, one tablespoon olive oil, one ounce feta, one cup plain kefir, one celery stalk, one tablespoon pumpkin seeds, and one cup chamomile tea." (difficulty β€”)2.8s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.8s
1 Β· TTS said I need to resolve one tablespoon pumpkin seeds before I log this meal. What should I use for one tablespoon pumpkin seeds?
2 Β· Card shown I need to resolve one tablespoon pumpkin seeds before I log this meal. What should I use for one tablespoon pumpkin seeds?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:52:06.625Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Log chicken and one medium Granny Smith apple." (difficulty β€”)0.5s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 Β· TTS said Logged Chicken breast and Apple. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Chicken breast and Apple. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Chicken breast Γ—1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F
created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:52:18.321Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had ice cream and one pint Halo Top sea salt caramel." (difficulty β€”)2.2s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.2s
1 Β· TTS said I need to resolve one pint Halo Top sea salt caramel before I log this meal. What should I use for one pint Halo Top sea salt caramel?
2 Β· Card shown I need to resolve one pint Halo Top sea salt caramel before I log this meal. What should I use for one pint Halo Top sea salt caramel?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:52:31.625Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Snack was a protein bar, one cup raspberries, and one mozzarella stick." (difficulty β€”)0.3s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.3s
1 Β· TTS said Logged a protein bar, one cup raspberries, and one mozzarella stick. Assumed 1 protein bar as the default serving. Tell me if you want a different amount. Tell me if that is not right.
2 Β· Card shown Logged a protein bar, one cup raspberries, and one mozzarella stick. Assumed 1 protein bar as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Generic protein bar Γ—1 (1 protein bar) 200 cal Β· 20g P Β· 22g C Β· 7g F
created food_log_entry: Raspberries Γ—1 (1 cup) 64 cal Β· 1.5g P Β· 14.6g C Β· 0.9g F
created food_log_entry: Mozzarella stick Γ—1 (1 mozzarella stick) 91 cal Β· 4.2g P Β· 7g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:52:43.019Z
6 Β· Why this food
  • The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • Matched the common-food list to "Generic protein bar" at 90% confidence.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 1.3 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I ate some chips, one Subway turkey six inch on wheat, and a bottle of water." (difficulty β€”)1.0s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS said Logged Potato chips, Subway 6" Oven-Roasted Turkey, and Water. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Potato chips, Subway 6" Oven-Roasted Turkey, and Water. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Potato chips Γ—1 (28 g) 150 cal Β· 2g P Β· 14.8g C Β· 9.8g F
created food_log_entry: Subway 6" Oven-Roasted Turkey Γ—1 (1 sandwich (6-inch default)) 480 cal Β· 26g P Β· 42g C Β· 23g F
created food_log_entry: Water Γ—1 (240 g) 0 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:52:55.216Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Lunch: one Chipotle burrito bowl with chicken black beans and fajita veggies, yogurt, one clementine, and one ounce almonds." (difficulty β€”)1.9s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.9s
1 Β· TTS said I need to resolve one Chipotle burrito bowl with chicken black beans and fajita veggies before I log this meal. What should I use for one Chipotle burrito bowl with chicken black beans and fajita veggies?
2 Β· Card shown I need to resolve one Chipotle burrito bowl with chicken black beans and fajita veggies before I log this meal. What should I use for one Chipotle burrito bowl with chicken black beans and fajita veggies?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:53:08.230Z
6 Β· Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 1.8 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had a McDonald's Big Mac, cereal, one cup skim milk, one kiwi, and one tablespoon peanut butter on celery." (difficulty β€”)1.3s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 Β· TTS said I need to resolve cereal before I log this meal. What should I use for cereal?
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:53:20.669Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Track toast, one cup Daisy 2 percent cottage cheese, one boiled egg, one cup cantaloupe, one ounce cashews, one Diet Dr Pepper, and one square Lindt 70 percent chocolate." (difficulty β€”)3.4s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.4s
1 Β· TTS said I need to resolve one ounce cashews before I log this meal. What should I use for one ounce cashews?
2 Β· Card shown I need to resolve one ounce cashews before I log this meal. What should I use for one ounce cashews?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:53:35.168Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould ask a clarifying question"Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot." (difficulty β€”)3.6s
Verdict Expected CLARIFY β€” should ask a clarifying question. FAIL: Logged a BLIND guess β€” no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β€” no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS said Logging your Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot β€” still looking up the nutrition details; they'll appear in your log shortly.
2 Β· Card shown Logging your Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot β€” still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot (pending)
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:53:57.035Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had peanut butter and one Fairlife chocolate shake." (difficulty β€”)0.6s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS said What exact ingredient should I use for one fairlife chocolate?
2 Β· Card shown What exact ingredient should I use for one fairlife chocolate?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:54:01.824Z
6 Β· Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 1.3 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Log a cup of yogurt, one slice Dave's Killer Bread 21 whole grains, and one cup green tea." (difficulty β€”)1.1s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.1s
1 Β· TTS said Logged Plain Greek yogurt, Bread, and one cup green tea. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Plain Greek yogurt, Bread, and one cup green tea. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Plain Greek yogurt Γ—1 (a cup (245 g)) 238 cal Β· 22.1g P Β· 9.6g C Β· 12.3g F
created food_log_entry: Bread Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
created food_log_entry: Green tea Γ—1 (1 cup) 2 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:54:14.138Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Afternoon: one Taco Bell crunchy taco supreme, hummus, one cup pineapple chunks, one tablespoon ranch on carrots, and one Siggi's vanilla skyr." (difficulty β€”)1.2s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.2s
1 Β· TTS said Logged Taco Bell Crunchy Taco Supreme, Hummus, one cup pineapple chunks, Baby carrots, and Skyr. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Taco Bell Crunchy Taco Supreme, Hummus, one cup pineapple chunks, Baby carrots, and Skyr. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Taco Bell Crunchy Taco Supreme Γ—1 (1 menu item) 190 cal Β· 8g P Β· 15g C Β· 11g F
created food_log_entry: Hummus Γ—1 (100 g) 166 cal Β· 7.9g P Β· 14.3g C Β· 9.6g F
created food_log_entry: Pineapple Γ—1 (1 cup) 83 cal Β· 0.8g P Β· 21.6g C Β· 0.2g F
created food_log_entry: Baby carrots Γ—1 (one tablespoon (9.3 g)) 3 cal Β· 0.1g P Β· 0.8g C Β· 0g F
created food_log_entry: Skyr Γ—1 (170 g) 107 cal Β· 18.7g P Β· 6.8g C Β· 0.3g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:54:26.587Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I ate a Panera broccoli cheddar soup bread bowl, one clementine, a protein shake, one cup cucumber slices, one ounce swiss cheese, one cup peppermint tea, and one madeleine cookie." (difficulty β€”)3.2s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.2s
1 Β· TTS said I need to resolve one madeleine cookie before I log this meal. What should I use for one madeleine cookie?
2 Β· Card shown I need to resolve one madeleine cookie before I log this meal. What should I use for one madeleine cookie?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:54:40.893Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Road trip: one bag Snyder's pretzels, one bottle Gatorade fruit punch, one ounce Jack Link's beef jerky, one Envy apple, one cheese stick, one cup grapes, and toast." (difficulty β€”)1.9s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.9s
1 Β· TTS said Logged Pretzels, Fruit Punch Thirst Quencher β€” GATORADE, Beef jerky, Apple, one cheese stick, one cup grapes, and Toast. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Pretzels, Fruit Punch Thirst Quencher β€” GATORADE, Beef jerky, Apple, one cheese stick, one cup grapes, and Toast. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Pretzels Γ—1 (28 g) 106 cal Β· 2.5g P Β· 22.1g C Β· 1g F
created food_log_entry: Fruit Punch Thirst Quencher Γ—1 (1 Bottle) 142 cal Β· 0g P Β· 36.1g C Β· 0g F
created food_log_entry: Beef jerky Γ—1 (one ounce (28.3 g)) 116 cal Β· 9.4g P Β· 3.1g C Β· 7.1g F
created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
created food_log_entry: Mozzarella stick Γ—1 (1 cheese stick) 91 cal Β· 4.2g P Β· 7g C Β· 5g F
created food_log_entry: Grapes Γ—1 (1 cup) 104 cal Β· 1.1g P Β· 27.3g C Β· 0.3g F
created food_log_entry: Toast Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:54:53.889Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"I drank one can La Croix pamplemousse and ate one pack SkinnyPop original." (difficulty β€”)2.8s
Verdict Expected LOG β€” should log the entry. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.8s
1 Β· TTS said I need to resolve one can La Croix pamplemousse before I log this meal. What should I use for one can La Croix pamplemousse?
2 Β· Card shown I need to resolve one can La Croix pamplemousse before I log this meal. What should I use for one can La Croix pamplemousse?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:55:08.113Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"I had one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios." (difficulty β€”)3.6s
Verdict Expected LOG β€” should log the entry. UNVERIFIED: UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS said Logging your one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios β€” still looking up the nutrition details; they'll appear in your log shortly.
2 Β· Card shown Logging your one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios β€” still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios (pending)
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:55:29.469Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Log one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar." (difficulty β€”)3.7s
Verdict Expected LOG β€” should log the entry. UNVERIFIED: UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS said Logging your one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar β€” still looking up the nutrition details; they'll appear in your log shortly.
2 Β· Card shown Logging your one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar β€” still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar (pending)
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:55:44.541Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"I ate one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream." (difficulty β€”)3.9s
Verdict Expected LOG β€” should log the entry. UNVERIFIED: UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.9s
1 Β· TTS said Logging your one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream β€” still looking up the nutrition details; they'll appear in your log shortly.
2 Β· Card shown Logging your one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream β€” still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream (pending)
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:55:59.169Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"Meal: one Impossible Whopper from Burger King, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard." (difficulty β€”)5.3s
Verdict Expected LOG β€” should log the entry. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 5.3s
1 Β· TTS said Which exact meal: one impossible whopper from burger king, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard should replace Pretzels?
2 Β· Card shown Which exact meal: one impossible whopper from burger king, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard should replace Pretzels?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:56:20.660Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"Track one cup overnight oats with chia, one scoop Vital Proteins collagen peptides, one cup soy milk, one tablespoon maple syrup, one medium blood orange, one ounce hemp hearts, and one cup plain kefir." (difficulty β€”)5.5s
Verdict Expected LOG β€” should log the entry. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 5.5s
1 Β· TTS said Which chia should I use?
2 Β· Card shown Which chia should I use?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:56:37.394Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig." (difficulty β€”)3.7s
Verdict Expected LOG β€” should log the entry. UNVERIFIED: UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β€” no saved row captured (untagged / lossy) β€” resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS said Logging your Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig β€” still looking up the nutrition details; they'll appear in your log shortly.
2 Β· Card shown Logging your Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig β€” still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig (pending)
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:57:00.880Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Log curry and rice with a side Caesar salad." (difficulty β€”)1.7s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.7s
1 Β· TTS said I need to resolve curry before I log this meal. What should I use for curry?
2 Β· Card shown I need to resolve curry before I log this meal. What should I use for curry?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:57:05.080Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had a smoothie, a salad, and soup." (difficulty β€”)2.2s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.2s
1 Β· TTS said I need to resolve a smoothie and a salad before I log this meal. What should I use for a smoothie and a salad?
2 Β· Card shown I need to resolve a smoothie and a salad before I log this meal. What should I use for a smoothie and a salad?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:57:18.383Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Dinner was steak, pasta, wine, bread, and dessert." (difficulty β€”)0.1s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.1s
1 Β· TTS said I did not log that yet because the exact item and amount decide the nutrition. What food and how much should I log?
2 Β· Card shown I did not log that yet because the exact item and amount decide the nutrition. What food and how much should I log?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:57:29.575Z
6 Β· Why this food
  • The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • Matched your saved foods and aliases to "bread" at 90% confidence.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 1.4 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had toast and chips." (difficulty β€”)0.4s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 Β· TTS said Logged Toast and Potato chips. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Toast and Potato chips. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Toast Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
created food_log_entry: Potato chips Γ—1 (28 g) 150 cal Β· 2g P Β· 14.8g C Β· 9.8g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:57:41.069Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Log chicken and yogurt." (difficulty β€”)0.9s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Why verdict Decisive guess with stated assumption + invited correction (good-guessing).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.9s
1 Β· TTS said Logged Chicken breast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
2 Β· Card shown Logged Chicken breast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Chicken breast Γ—1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F
created food_log_entry: Plain Greek yogurt Γ—1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
created food_log_entry: Chicken breast Γ—1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F
created food_log_entry: Plain Greek yogurt Γ—1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:57:53.116Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould ask a clarifying question"I ate a sandwich, some soup, a cookie, milk, and fruit." (difficulty β€”)2.9s
Verdict Expected CLARIFY β€” should ask a clarifying question. FAIL: Logged a BLIND guess β€” no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β€” no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.9s
1 Β· TTS said I need to resolve a sandwich, a cookie, and fruit before I log this meal. What should I use for a sandwich, a cookie, and fruit?
2 Β· Card shown I need to resolve a sandwich, a cookie, and fruit before I log this meal. What should I use for a sandwich, a cookie, and fruit?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:58:07.158Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had a big salad with chicken." (difficulty β€”)2.3s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.3s
1 Β· TTS said Which type of big salad with chicken was it? There are several with different nutrition.
2 Β· Card shown Which type of big salad with chicken was it? There are several with different nutrition.
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:58:20.556Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould ask a clarifying question"Breakfast: two eggs, toast, bacon, coffee, oatmeal, orange juice, and pancakes." (difficulty β€”)1.6s
Verdict Expected CLARIFY β€” should ask a clarifying question. FAIL: Logged a BLIND guess β€” no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β€” no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.6s
1 Β· TTS said I did not make any app changes. What would you like me to do with that?
2 Β· Card shown I did not make any app changes. What would you like me to do with that?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:58:33.314Z
6 Β· Why this food
  • The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • Matched your saved foods and aliases to "orange" at 90% confidence.
  • Matched your past food logs to "Toast" at 82% confidence.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 1.8 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Buffet plate: chicken, rice, salad, soup, bread, pasta, fish, vegetables, dessert, and coffee." (difficulty β€”)1.4s
Verdict Expected LOG β€” should log the entry. UNVERIFIED: Informational non-gating row (excluded from pass/fail tallies).
Why verdict Informational non-gating row (excluded from pass/fail tallies).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.4s
1 Β· TTS said I’ll log your buffet plate β€” still looking up the nutrition details; they’ll appear in your log shortly.
2 Β· Card shown I’ll log your buffet plate β€” still looking up the nutrition details; they’ll appear in your log shortly.
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:58:45.875Z
6 Β· Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 2.5 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould ask a clarifying question"I had a ham and swiss on rye and some chips." (difficulty β€”)2.0s
Verdict Expected CLARIFY β€” should ask a clarifying question. FAIL: Logged a BLIND guess β€” no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β€” no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.0s
1 Β· TTS said I need to resolve rye before I log this meal. What should I use for rye?
2 Β· Card shown I need to resolve rye before I log this meal. What should I use for rye?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:58:59.004Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Log curry and rice." (difficulty β€”)2.0s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.0s
1 Β· TTS said I need to resolve curry before I log this meal. What should I use for curry?
2 Β· Card shown I need to resolve curry before I log this meal. What should I use for curry?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:59:12.108Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"Morning: toast, cereal, yogurt, a protein shake, two eggs, one banana, and coffee." (difficulty β€”)1.5s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.5s
1 Β· TTS said I did not make any app changes. What would you like me to do with that?
2 Β· Card shown I did not make any app changes. What would you like me to do with that?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:59:24.685Z
6 Β· Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 Β· Lookup method
  • The report captured the app food lookup that ran before or alongside the final decision.
  • It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
  • The lookup evidence took 2.0 seconds to collect.
  • This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana." (difficulty β€”)3.7s
Verdict Expected LOG β€” should log the entry. UNVERIFIED: Informational non-gating row (excluded from pass/fail tallies).
Why verdict Informational non-gating row (excluded from pass/fail tallies).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS said Logging your Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana β€” still looking up the nutrition details; they'll appear in your log shortly.
2 Β· Card shown Logging your Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana β€” still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana (pending)
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:59:39.572Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould ask a clarifying question"I had pizza, cereal, yogurt, hummus, and one Coca-Cola Zero." (difficulty β€”)1.5s
Verdict Expected CLARIFY β€” should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.5s
1 Β· TTS said I need to resolve cereal before I log this meal. What should I use for cereal?
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T08:59:52.320Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould ask a clarifying question"I had a smoothie for breakfast, two eggs, and toast." (difficulty β€”)2.7s
Verdict Expected CLARIFY β€” should ask a clarifying question. FAIL: Logged a BLIND guess β€” no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β€” no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.7s
1 Β· TTS said I need to resolve a smoothie before I log this meal. What should I use for a smoothie?
2 Β· Card shown I need to resolve a smoothie before I log this meal. What should I use for a smoothie?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T09:00:06.188Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the entry"Snack was one cup red seedless grapes, one ounce manchego cheese, and one Wasa multigrain crispbread." (difficulty β€”)1.6s
Verdict Expected LOG β€” should log the entry. PASS: Logged (write-truth verified): Multigrain Crispbread, Grapes, Manchego cheese β€” card not captured.
Why verdict Logged (write-truth verified): Multigrain Crispbread, Grapes, Manchego cheese β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.6s
1 Β· TTS said Logged Grapes, one ounce manchego cheese, and Multigrain Crispbread β€” DIVINA. Assumed 1 oz as the default serving. Tell me if you want a different amount. Tell me if that is not right.
2 Β· Card shown Logged Grapes, one ounce manchego cheese, and Multigrain Crispbread β€” DIVINA. Assumed 1 oz as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Grapes Γ—1 (one cup (151 g)) 104 cal Β· 1.1g P Β· 27.3g C Β· 0.3g F
created food_log_entry: Manchego cheese Γ—1 (1 oz) 113 cal Β· 7.1g P Β· 0.1g C Β· 9.4g F
created food_log_entry: Multigrain Crispbread Γ—1 (1 PIECE) 120 cal Β· 4g P Β· 13g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T09:00:19.131Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the entry"Lunch: one Sweetgreen harvest bowl, one cup miso soup, one sheet nori, one tablespoon sesame seeds, and one cup jasmine green tea." (difficulty β€”)1.0s
Verdict Expected LOG β€” should log the entry. FAIL: WRITE-TRUTH FAIL β€” WRONG/MISSING item "nori sheet" β€” no saved row with matching identity (rows: Sweetgreen Harvest Bowl, Miso soup); WRONG/MISSING item "one tablespoon sesame seeds" β€” no saved row with matching identity (rows: Sweetgreen Harvest Bowl, Miso soup); WRONG/MISSING item "one cup jasmine green tea" β€” no saved row with matching identity (rows: Sweetgreen Harvest Bowl, Miso soup)
Why verdict WRITE-TRUTH FAIL β€” WRONG/MISSING item "nori sheet" β€” no saved row with matching identity (rows: Sweetgreen Harvest Bowl, Miso soup); WRONG/MISSING item "one tablespoon sesame seeds" β€” no saved row with matching identity (rows: Sweetgreen Harvest Bowl, Miso soup); WRONG/MISSING item "one cup jasmine green tea" β€” no saved row with matching identity (rows: Sweetgreen Harvest Bowl, Miso soup)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS said Logged Sweetgreen Harvest Bowl, one cup miso soup, one sheet nori, one tablespoon sesame seeds, and one cup jasmine green tea. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
2 Β· Card shown Logged Sweetgreen Harvest Bowl, one cup miso soup, one sheet nori, one tablespoon sesame seeds, and one cup jasmine green tea. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Sweetgreen Harvest Bowl Γ—1 (bowl) 540 cal Β· 28g P Β· 48g C Β· 24g F
created food_log_entry: Miso soup Γ—1 (1 cup) 49 cal Β· 3.7g P Β· 6.6g C Β· 1.5g F
created food_log_entry: Nori Γ—1 (1 sheet nori) 8 cal Β· 1.2g P Β· 1.1g C Β· 0.1g F
created food_log_entry: Sesame seeds Γ—1 (1 tbsp) 52 cal Β· 1.6g P Β· 2.1g C Β· 4.5g F
created food_log_entry: Green tea Γ—1 (1 cup) 2 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T09:00:31.251Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)