multifood-corpus-a Β· 1 hour 17 minutes ago Β· iOS sim
Rows
60
Pass
42 (70%)
Fail
12 (20%)
Unverified
6 (10%)
Pass rate
78%
Avg difficulty
β
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace β needs a human look (that's you π/π-ing it). "Pass rate" = pass Γ· (pass + fail) β it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows). Unverified breaks down as: 6 unclassified β the "grader coverage gap" ones aren't ambiguous, the grader just doesn't fully score that intent type yet.
Why the fails happened β comprehension vs execution vs cosmetic
Comprehension β picked the wrong action/target (the hard problem)
11 (100%)
resolution β 11 (100% of comprehension)
"Log a medium Fuji apple and one KIND peanut butter dark chocolate bar." β Logged the WRONG or a MISSING item vs what was asked (saved row identity mismatch).
"I had a Chipotle chicken burrito bowl and a bottle of Topo Chico." β Over-asked: asked instead of logging (no saved row).
"Breakfast was two scrambled eggs, one slice of sourdough toast, and eight ounces of cold brew." β Over-asked: asked instead of logging (no saved row).
"Track one Oikos triple zero vanilla yogurt, a medium navel orange, and twelve almonds." β Logged the WRONG or a MISSING item vs what was asked (saved row identity mismatch).
"I ate a Panera Mediterranean bowl, one Perfect Bar almond butter, a cup of bone broth, two celery stalks, and one Medjool date." β Logged the wrong AMOUNT vs the utterance (saved serving mismatch).
"Lunch was a Chipotle steak burrito, granola, one Fuji apple, one cup sparkling water, and a string cheese." β Under-asked: logged when it should have asked a clarifying question.
+ 5 more
Of 11 fails: if most are comprehension, that's the hard problem (the app didn't figure out the right action/target); execution/cosmetic means the decision was right but something downstream broke. Click a bar for the sub-split + example utterances.
Handled correctly? β by expected action
Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed to
N
Correct
Wrong
Unverified
βΈ CLARIFY β ask a clarifying question
39
36 (92%)
3 (8%)
0 (0%)
3 handled wrong β click one to jump to its full detail below
"Lunch was a Chipotle steak burrito, granola, one Fuji apple, one cup sparkling water, and a string cheese."
Logged a BLIND guess β no stated assumption, no correction invited.
"Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent."
Logged a BLIND guess β no stated assumption, no correction invited.
"Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot."
Logged a BLIND guess β no stated assumption, no correction invited.
βΈ LOG β log the entry
21
6 (29%)
9 (43%)
6 (29%)
9 handled wrong β click one to jump to its full detail below
"Log a medium Fuji apple and one KIND peanut butter dark chocolate bar."
WRITE-TRUTH FAIL β WRONG/MISSING item "KIND peanut butter dark chocolate bar" β no saved row with matching identity (rows: Apple, Chocolate Candy, Peanut Butter Filled)
"I had a Chipotle chicken burrito bowl and a bottle of Topo Chico."
OVER-ASK β asked instead of logging (no saved row).
"Breakfast was two scrambled eggs, one slice of sourdough toast, and eight ounces of cold brew."
OVER-ASK β asked instead of logging (no saved row).
"Track one Oikos triple zero vanilla yogurt, a medium navel orange, and twelve almonds."
WRITE-TRUTH FAIL β WRONG/MISSING item "Oikos triple zero vanilla yogurt" β no saved row with matching identity (rows: Plain Greek yogurt, Orange, Almonds)
"I ate a Panera Mediterranean bowl, one Perfect Bar almond butter, a cup of bone broth, two celery stalks, and one Medjool date."
WRITE-TRUTH FAIL β IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving])
"I drank one can La Croix pamplemousse and ate one pack SkinnyPop original."
OVER-ASK β asked instead of logging (no saved row).
"Meal: one Impossible Whopper from Burger King, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard."
OVER-ASK β asked instead of logging (no saved row).
"Track one cup overnight oats with chia, one scoop Vital Proteins collagen peptides, one cup soy milk, one tablespoon maple syrup, one medium blood orange, one ounce hemp hearts, and one cup plain kefir."
OVER-ASK β asked instead of logging (no saved row).
OVER-ASK β asked instead of logging (no saved row).
Total
60
42 (78%)
12
6
Accuracy by difficulty
Pending A1's per-utterance difficulty score (requested 2026-07-05) β this bar chart lights up once that lands.
Clarification follow-ups β scored separately
Second turn: app asked, we replied β did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.
Cosmetic only
Not yet classified β pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.
System / infra
Not yet classified β pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.
Latency
Avg (time to ready)
2.2s
p90
3.7s
Max
5.7s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers β click any dot to jump to its detail.
0s
1s
2s
5s
6s
Response path β quick (single response) vs async (an ack like "Working on itβ¦" before the real answer).
Sync clarification
39
Quick response
21
Slowest 8 utterances (click to jump to detail):
"Track one cup overnight oats with chia, one scoop Vital Proteins collagen peptides, one cup soy milk, one tablespoon maple syrup, one medium blood orange, one ounce hemp hearts, and one cup plain kefir."5.7s
"Meal: one Impossible Whopper from Burger King, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard."5.3s
"I ate one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream."3.9s
"Log one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar."3.8s
"Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig."3.7s
"Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana."3.7s
"I ate pizza, one cup Greek salad with feta, one bottle La Croix lime, two eggs, one cup vanilla bean ice cream, one slice whole grain bread, and six ounces Atlantic cod."3.7s
"Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice."3.7s
Filter β controls the list below
Pass / Fail / Unverified
PASS 42FAIL 12UNVERIFIED 6
Module (intended for)
food (60)
Utterance sub-type (within module)
60 shown β 42 pass, 12 fail, 6 unverified
Per-utterance detail
FAILshould log the entry"Log a medium Fuji apple and one KIND peanut butter dark chocolate bar." (difficulty β)0.8s
Verdict Expected LOG β should log the entry. FAIL: WRITE-TRUTH FAIL β WRONG/MISSING item "KIND peanut butter dark chocolate bar" β no saved row with matching identity (rows: Apple, Chocolate Candy, Peanut Butter Filled)
Why verdict WRITE-TRUTH FAIL β WRONG/MISSING item "KIND peanut butter dark chocolate bar" β no saved row with matching identity (rows: Apple, Chocolate Candy, Peanut Butter Filled)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.8s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Apple and Chocolate Candy, Peanut Butter Filled. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Apple Γ1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F created food_log_entry: Chocolate Candy, Peanut Butter Filled Γ1 (1 large/king size) 438 cal Β· 7g P Β· 57.7g C Β· 22.2g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:55:41.017Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould log the entry"I had a Chipotle chicken burrito bowl and a bottle of Topo Chico." (difficulty β)3.1s
Verdict Expected LOG β should log the entry. FAIL: OVER-ASK β asked instead of logging (no saved row).
Why verdict OVER-ASK β asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.1s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve a Chipotle chicken burrito bowl before I log this meal. What should I use for a Chipotle chicken burrito bowl?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:55:55.234Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould log the entry"Breakfast was two scrambled eggs, one slice of sourdough toast, and eight ounces of cold brew." (difficulty β)3.6s
Verdict Expected LOG β should log the entry. FAIL: OVER-ASK β asked instead of logging (no saved row).
Why verdict OVER-ASK β asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve eight ounces of cold brew before I log this meal. What should I use for eight ounces of cold brew?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:56:10.052Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould log the entry"Track one Oikos triple zero vanilla yogurt, a medium navel orange, and twelve almonds." (difficulty β)0.7s
Verdict Expected LOG β should log the entry. FAIL: WRITE-TRUTH FAIL β WRONG/MISSING item "Oikos triple zero vanilla yogurt" β no saved row with matching identity (rows: Plain Greek yogurt, Orange, Almonds)
Why verdict WRITE-TRUTH FAIL β WRONG/MISSING item "Oikos triple zero vanilla yogurt" β no saved row with matching identity (rows: Plain Greek yogurt, Orange, Almonds)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.7s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Plain Greek yogurt, Orange, and twelve almonds. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Plain Greek yogurt Γ1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F created food_log_entry: Orange Γ1 (184 g) 86 cal Β· 1.7g P Β· 21.7g C Β· 0.2g F created food_log_entry: Almonds Γ1 (12 almonds) 83 cal Β· 3.1g P Β· 3.1g C Β· 7.2g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:56:21.949Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould log the entry"Lunch was six ounces grilled Atlantic salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge." (difficulty β)0.6s
Verdict Expected LOG β should log the entry. PASS: Logged (write-truth verified): Olive oil, Broccoli, Cooked brown rice, Lemon, Salmon β card not captured.
Why verdict Logged (write-truth verified): Olive oil, Broccoli, Cooked brown rice, Lemon, Salmon β card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Salmon Γ1 (six ounces (170.1 g)) 354 cal Β· 34g P Β· 0g C Β· 22.1g F created food_log_entry: Broccoli Γ1 (1 cup) 55 cal Β· 3.8g P Β· 11.3g C Β· 0.6g F created food_log_entry: Cooked brown rice Γ1 (0.5 cup) 109 cal Β· 2.2g P Β· 22.9g C Β· 0.8g F created food_log_entry: Olive oil Γ1 (1 tbsp) 120 cal Β· 0g P Β· 0g C Β· 13.6g F created food_log_entry: Lemon Γ1 (1 lemon wedge) 17 cal Β· 0.6g P Β· 5.4g C Β· 0.2g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:56:33.669Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould log the entry"I ate a Panera Mediterranean bowl, one Perfect Bar almond butter, a cup of bone broth, two celery stalks, and one Medjool date." (difficulty β)1.1s
Verdict Expected LOG β should log the entry. FAIL: WRITE-TRUTH FAIL β IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving])
Why verdict WRITE-TRUTH FAIL β IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving])
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.1s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Panera Bread Mediterranean Bowl, Almond Butter β PERFECT BAR, a cup of bone broth, two celery stalks, and one medjool date. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Panera Bread Mediterranean Bowl Γ1 (1 Bowl) 590 cal Β· 16g P Β· 66g C Β· 31g F created food_log_entry: Almond Butter Γ1 (1 BAR) 320 cal Β· 13g P Β· 25g C Β· 19g F created food_log_entry: Bone broth Γ1 (1 cup) 36 cal Β· 7.2g P Β· 1.2g C Β· 0.7g F created food_log_entry: Celery Γ1 (2 celery stalks) 13 cal Β· 0.6g P Β· 2.4g C Β· 0.2g F created food_log_entry: Dates Γ1 (1 medjool date) 66 cal Β· 0.4g P Β· 18g C Β· 0g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:56:45.918Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from restaurant menu.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould log the entry"Dinner: eight ounces lean ground turkey, one cup roasted Brussels sprouts, half avocado, one Thomas whole wheat English muffin, one Laughing Cow wedge, ten baby carrots, and sixteen ounces sparkling water." (difficulty β)1.3s
Verdict Expected LOG β should log the entry. PASS: Logged (write-truth verified): Brussels sprouts, Baby carrots, Lean ground turkey, Laughing Cow, Avocado, Water, Whole wheat English muffin β card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged eight ounces lean ground turkey, one cup roasted brussels sprouts, half avocado, Whole wheat English muffin, one laughing cow wedge, ten baby carrots, and Water. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Lean ground turkey Γ1 (8 oz) 386 cal Β· 61.2g P Β· 0g C Β· 15.9g F created food_log_entry: Brussels sprouts Γ1 (1 cup) 70 cal Β· 5.4g P Β· 14.4g C Β· 0.5g F created food_log_entry: Avocado Γ1 (half avocado) 120 cal Β· 1.5g P Β· 6.4g C Β· 11g F created food_log_entry: Whole wheat English muffin Γ1 (1 muffin (57 g)) 128 cal Β· 5.6g P Β· 24.8g C Β· 1.3g F created food_log_entry: Laughing Cow Γ1 (1 laughing cow wedge) 35 cal Β· 2g P Β· 1g C Β· 2.5g F created food_log_entry: Baby carrots Γ1 (10 baby carrots) 35 cal Β· 0.6g P Β· 8.2g C Β· 0.1g F created food_log_entry: Water Γ1 (sixteen ounces (453.6 g)) 0 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:56:58.367Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould log the entry"Log a Starbucks tall oat milk latte, one Kodiak Cakes blueberry waffle, two strips turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and one Justin maple almond butter squeeze pack." (difficulty β)1.3s
Verdict Expected LOG β should log the entry. PASS: Logged (write-truth verified): Blueberry waffle, Oat milk latte, Egg, Almond butter, Mozzarella, Turkey bacon, Blackberries β card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged a starbucks tall oat milk latte, one kodiak cakes blueberry waffle, Turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and Almond butter. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Oat milk latte Γ1 (1 starbucks tall oat milk latte) 135 cal Β· 4.3g P Β· 18.5g C Β· 5.3g F created food_log_entry: Blueberry waffle Γ1 (1 kodiak cakes blueberry waffle) 180 cal Β· 7g P Β· 24.5g C Β· 4.9g F created food_log_entry: Turkey bacon Γ2 (14 g) 60 cal Β· 8.2g P Β· 1g C Β· 3g F created food_log_entry: Blackberries Γ1 (1 cup) 62 cal Β· 2g P Β· 14.7g C Β· 0.7g F created food_log_entry: Egg Γ1 (1 egg) 72 cal Β· 6.3g P Β· 0.4g C Β· 4.8g F created food_log_entry: Mozzarella Γ1 (1 oz) 72 cal Β· 6.9g P Β· 0.8g C Β· 4.5g F created food_log_entry: Almond butter Γ1 (32 g) 196 cal Β· 6.7g P Β· 6.1g C Β· 17.9g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:57:10.801Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould log the entry"Snack spread: one Blue Diamond smokehouse almonds pack, one Babybel light, one cup edamame, one Pink Lady apple, one cup coconut water, one Good Culture strawberry cottage cheese cup, one rice cake, one tablespoon chia seeds, five walnuts, and one Clif Kid Zbar." (difficulty β)2.0s
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.0s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Almonds, one babybel light, one cup edamame, Apple, one cup coconut water, Cottage cheese, one rice cake, one tablespoon chia seeds, five walnuts, and one clif kid zbar. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Almonds Γ1 (28 g) 162 cal Β· 5.9g P Β· 6g C Β· 14g F created food_log_entry: Babybel light Γ1 (1 babybel light) 42 cal Β· 6g P Β· 0g C Β· 2.5g F created food_log_entry: Edamame Γ1 (1 cup) 188 cal Β· 18.4g P Β· 13.8g C Β· 8.1g F created food_log_entry: Apple Γ1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F created food_log_entry: Coconut water Γ1 (1 cup) 46 cal Β· 1.7g P Β· 8.9g C Β· 0.5g F created food_log_entry: Cottage cheese Γ1 (1 cup) 221 cal Β· 25.1g P Β· 7.7g C Β· 9.7g F created food_log_entry: Plain rice cakes Γ1 (1 rice cake) 35 cal Β· 0.7g P Β· 7.3g C Β· 0.3g F created food_log_entry: Chia seeds Γ1 (1 tbsp) 51 cal Β· 1.7g P Β· 4.4g C Β· 3.2g F created food_log_entry: Walnuts Γ1 (5 walnuts) 65 cal Β· 1.5g P Β· 1.4g C Β· 6.5g F created food_log_entry: Clif Kid Zbar Γ1 (1 clif kid zbar) 130 cal Β· 3g P Β· 23g C Β· 4g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:57:24.167Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice." (difficulty β)3.7s
Verdict Expected LOG β should log the entry. UNVERIFIED: UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:57:45.700Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had toast and a Chobani peach Greek yogurt." (difficulty β)0.4s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Toast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Toast Γ1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F created food_log_entry: Plain Greek yogurt Γ1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:57:50.468Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Log cottage cheese and one large Honeycrisp apple." (difficulty β)0.6s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Cottage cheese and Apple. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Cottage cheese Γ1 (100 g) 98 cal Β· 11.1g P Β· 3.4g C Β· 4.3g F created food_log_entry: Apple Γ1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:58:02.301Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I ate a turkey sandwich on rye, some chips, and a Diet Coke." (difficulty β)3.2s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.2s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve a turkey sandwich on rye before I log this meal. What should I use for a turkey sandwich on rye?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:58:16.636Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Breakfast was two eggs, toast, and black coffee." (difficulty β)0.1s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.1s
1 Β· TTS saidno speech captured
2 Β· Card shown I did not log that yet because the exact item and amount decide the nutrition. What food and how much should I log?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:58:27.834Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had a protein shake, one RXBAR chocolate sea salt bar, and a banana." (difficulty β)0.9s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.9s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Protein powder, Chocolate Sea Salt Bar β RXBAR, and a banana. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Protein powder Γ1 (100 g) 367 cal Β· 83g P Β· 3.3g C Β· 1.7g F created food_log_entry: Chocolate Sea Salt Bar Γ1 (1 bar) 200 cal Β· 12g P Β· 23g C Β· 8g F created food_log_entry: Banana Γ1 (1 banana) 105 cal Β· 1.3g P Β· 27.1g C Β· 0.4g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:58:39.901Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould ask a clarifying question"Lunch was a Chipotle steak burrito, granola, one Fuji apple, one cup sparkling water, and a string cheese." (difficulty β)3.6s
Verdict Expected CLARIFY β should ask a clarifying question. FAIL: Logged a BLIND guess β no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your Lunch was a Chipotle steak burrito, granola, one Fuji apple, one cup sparkling water, and a string cheese β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Lunch was a Chipotle steak burrito, granola, one Fuji apple, one cup sparkling water, and a string cheese (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:59:01.618Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had six ounces grilled salmon, peanut butter, one cup steamed asparagus, one Diet Pepsi, and one ounce sharp cheddar." (difficulty β)0.7s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.7s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged six ounces grilled salmon, Peanut butter, one cup steamed asparagus, Diet cola, and Cheddar cheese. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Salmon Γ1 (6 oz) 354 cal Β· 34g P Β· 0g C Β· 22.1g F created food_log_entry: Peanut butter Γ1 (32 g) 188 cal Β· 8g P Β· 6.3g C Β· 16.1g F created food_log_entry: Asparagus Γ1 (1 cup) 40 cal Β· 4.3g P Β· 7.4g C Β· 0.4g F created food_log_entry: Diet cola Γ1 (355 g) 0 cal Β· 0g P Β· 0g C Β· 0g F created food_log_entry: Cheddar cheese Γ1 (one ounce (28.3 g)) 114 cal Β· 6.5g P Β· 1g C Β· 9.4g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:59:06.520Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Track a Wendy's spicy chicken sandwich, cereal, one cup two percent milk, one clementine, and one ounce provolone." (difficulty β)1.9s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.9s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:59:19.534Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould ask a clarifying question"Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent." (difficulty β)3.6s
Verdict Expected CLARIFY β should ask a clarifying question. FAIL: Logged a BLIND guess β no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:59:41.184Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I ate pizza, one cup Greek salad with feta, one bottle La Croix lime, two eggs, one cup vanilla bean ice cream, one slice whole grain bread, and six ounces Atlantic cod." (difficulty β)3.7s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve one cup vanilla bean ice cream before I log this meal. What should I use for one cup vanilla bean ice cream?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T11:59:49.067Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Log yogurt, one medium Bartlett pear, one ounce pistachios, one In-N-Out grilled cheese, one tablespoon Jif creamy peanut butter, one cup peppermint tea, and one mandarin." (difficulty β)3.3s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.3s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve one cup peppermint tea and one mandarin before I log this meal. What should I use for one cup peppermint tea and one mandarin?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:00:15.036Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Big breakfast: two eggs, one strip turkey bacon, one medium banana, one cup black coffee, one slice whole wheat bread, six ounces plain nonfat Greek yogurt, and toast." (difficulty β)1.2s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.2s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Turkey bacon. Using your recent Turkey bacon history. Tell me if that is wrong.
3 Β· App data rows written created food_log_entry: Turkey bacon Γ0.14285714285714285 (14 g) 4 cal Β· 0.6g P Β· 0.1g C Β· 0.2g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:00:27.368Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from user default.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Meal prep: four ounces chicken breast, one cup jasmine rice, one cup green beans, hummus, one tablespoon olive oil, one ounce feta, one cup plain kefir, one celery stalk, one tablespoon pumpkin seeds, and one cup chamomile tea." (difficulty β)2.9s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.9s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve one tablespoon pumpkin seeds before I log this meal. What should I use for one tablespoon pumpkin seeds?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:00:41.386Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Log chicken and one medium Granny Smith apple." (difficulty β)0.5s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Chicken breast and Apple. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Chicken breast Γ1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F created food_log_entry: Apple Γ1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:00:53.034Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had ice cream and one pint Halo Top sea salt caramel." (difficulty β)2.7s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.7s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve one pint Halo Top sea salt caramel before I log this meal. What should I use for one pint Halo Top sea salt caramel?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:01:06.953Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Snack was a protein bar, one cup raspberries, and one mozzarella stick." (difficulty β)0.3s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.3s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged a protein bar, one cup raspberries, and one mozzarella stick. Assumed 1 protein bar as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Generic protein bar Γ1 (1 protein bar) 200 cal Β· 20g P Β· 22g C Β· 7g F created food_log_entry: Raspberries Γ1 (1 cup) 64 cal Β· 1.5g P Β· 14.6g C Β· 0.9g F created food_log_entry: Mozzarella stick Γ1 (1 mozzarella stick) 91 cal Β· 4.2g P Β· 7g C Β· 5g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:01:18.333Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I ate some chips, one Subway turkey six inch on wheat, and a bottle of water." (difficulty β)1.0s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Potato chips, Subway Oven Roasted Turkey, and Water. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Potato chips Γ1 (28 g) 150 cal Β· 2g P Β· 14.8g C Β· 9.8g F created food_log_entry: Subway Oven Roasted Turkey Γ1 (57 g) 60 cal Β· 11g P Β· 0g C Β· 1g F created food_log_entry: Water Γ1 (240 g) 0 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:01:30.418Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Lunch: one Chipotle burrito bowl with chicken black beans and fajita veggies, yogurt, one clementine, and one ounce almonds." (difficulty β)2.1s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.1s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve one Chipotle burrito bowl with chicken, black beans, and fajita veggies before I log this meal. What should I use for one Chipotle burrito bowl?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:01:43.635Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had a McDonald's Big Mac, cereal, one cup skim milk, one kiwi, and one tablespoon peanut butter on celery." (difficulty β)1.8s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.8s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:01:56.550Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Track toast, one cup Daisy 2 percent cottage cheese, one boiled egg, one cup cantaloupe, one ounce cashews, one Diet Dr Pepper, and one square Lindt 70 percent chocolate." (difficulty β)3.3s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.3s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve one ounce cashews before I log this meal. What should I use for one ounce cashews?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:02:11.017Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould ask a clarifying question"Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot." (difficulty β)3.6s
Verdict Expected CLARIFY β should ask a clarifying question. FAIL: Logged a BLIND guess β no stated assumption, no correction invited.
Why verdict Logged a BLIND guess β no stated assumption, no correction invited.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:02:32.715Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had peanut butter and one Fairlife chocolate shake." (difficulty β)0.6s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS saidno speech captured
2 Β· Card shown What exact ingredient should I use for one fairlife chocolate?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:02:37.452Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Log a cup of yogurt, one slice Dave's Killer Bread 21 whole grains, and one cup green tea." (difficulty β)1.1s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.1s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Plain Greek yogurt, Bread, and one cup green tea. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Plain Greek yogurt Γ1 (a cup (245 g)) 238 cal Β· 22.1g P Β· 9.6g C Β· 12.3g F created food_log_entry: Bread Γ1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F created food_log_entry: Green tea Γ1 (1 cup) 2 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:02:49.768Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Afternoon: one Taco Bell crunchy taco supreme, hummus, one cup pineapple chunks, one tablespoon ranch on carrots, and one Siggi's vanilla skyr." (difficulty β)1.2s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.2s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Taco Bell Crunchy Taco Supreme, Hummus, one cup pineapple chunks, Baby carrots, and Skyr. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Taco Bell Crunchy Taco Supreme Γ1 (1 menu item) 190 cal Β· 8g P Β· 15g C Β· 11g F created food_log_entry: Hummus Γ1 (100 g) 166 cal Β· 7.9g P Β· 14.3g C Β· 9.6g F created food_log_entry: Pineapple Γ1 (1 cup) 83 cal Β· 0.8g P Β· 21.6g C Β· 0.2g F created food_log_entry: Baby carrots Γ1 (one tablespoon (9.3 g)) 3 cal Β· 0.1g P Β· 0.8g C Β· 0g F created food_log_entry: Skyr Γ1 (170 g) 107 cal Β· 18.7g P Β· 6.8g C Β· 0.3g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:03:02.085Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from restaurant menu.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I ate a Panera broccoli cheddar soup bread bowl, one clementine, a protein shake, one cup cucumber slices, one ounce swiss cheese, one cup peppermint tea, and one madeleine cookie." (difficulty β)3.0s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.0s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve one madeleine cookie before I log this meal. What should I use for one madeleine cookie?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:03:16.186Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Road trip: one bag Snyder's pretzels, one bottle Gatorade fruit punch, one ounce Jack Link's beef jerky, one Envy apple, one cheese stick, one cup grapes, and toast." (difficulty β)1.8s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.8s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Pretzels, Fruit Punch Thirst Quencher β GATORADE, Beef jerky, Apple, one cheese stick, one cup grapes, and Toast. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Pretzels Γ1 (28 g) 106 cal Β· 2.5g P Β· 22.1g C Β· 1g F created food_log_entry: Fruit Punch Thirst Quencher Γ1 (1 Bottle) 142 cal Β· 0g P Β· 36.1g C Β· 0g F created food_log_entry: Beef jerky Γ1 (one ounce (28.3 g)) 116 cal Β· 9.4g P Β· 3.1g C Β· 7.1g F created food_log_entry: Apple Γ1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F created food_log_entry: Mozzarella stick Γ1 (1 cheese stick) 91 cal Β· 4.2g P Β· 7g C Β· 5g F created food_log_entry: Grapes Γ1 (1 cup) 104 cal Β· 1.1g P Β· 27.3g C Β· 0.3g F created food_log_entry: Toast Γ1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:03:29.152Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould log the entry"I drank one can La Croix pamplemousse and ate one pack SkinnyPop original." (difficulty β)3.6s
Verdict Expected LOG β should log the entry. FAIL: OVER-ASK β asked instead of logging (no saved row).
Why verdict OVER-ASK β asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS saidno speech captured
2 Β· Card shown What should I use for one can La Croix pamplemousse? I did not log it yet because I could not match it safely.
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:03:43.867Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"I had one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios." (difficulty β)3.6s
Verdict Expected LOG β should log the entry. UNVERIFIED: UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:04:05.549Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Log one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar." (difficulty β)3.8s
Verdict Expected LOG β should log the entry. UNVERIFIED: UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.8s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:04:20.365Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"I ate one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream." (difficulty β)3.9s
Verdict Expected LOG β should log the entry. UNVERIFIED: UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.9s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:04:35.499Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould log the entry"Meal: one Impossible Whopper from Burger King, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard." (difficulty β)5.3s
Verdict Expected LOG β should log the entry. FAIL: OVER-ASK β asked instead of logging (no saved row).
Why verdict OVER-ASK β asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 5.3s
1 Β· TTS saidno speech captured
2 Β· Card shown Which exact meal: one impossible whopper from burger king, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard should replace Grapes?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:04:56.336Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould log the entry"Track one cup overnight oats with chia, one scoop Vital Proteins collagen peptides, one cup soy milk, one tablespoon maple syrup, one medium blood orange, one ounce hemp hearts, and one cup plain kefir." (difficulty β)5.7s
Verdict Expected LOG β should log the entry. FAIL: OVER-ASK β asked instead of logging (no saved row).
Why verdict OVER-ASK β asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 5.7s
1 Β· TTS saidno speech captured
2 Β· Card shown Which chia should I use?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:05:13.150Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig." (difficulty β)3.7s
Verdict Expected LOG β should log the entry. UNVERIFIED: UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:05:35.049Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Log curry and rice with a side Caesar salad." (difficulty β)1.5s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.5s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve curry before I log this meal. What should I use for curry?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:05:40.600Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had a smoothie, a salad, and soup." (difficulty β)2.2s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.2s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve a smoothie and a salad before I log this meal. What should I use for a smoothie and a salad?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:05:53.967Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Dinner was steak, pasta, wine, bread, and dessert." (difficulty β)0.1s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.1s
1 Β· TTS saidno speech captured
2 Β· Card shown I did not log that yet because the exact item and amount decide the nutrition. What food and how much should I log?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:06:05.134Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had toast and chips." (difficulty β)0.4s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Toast and Potato chips. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Toast Γ1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F created food_log_entry: Potato chips Γ1 (28 g) 150 cal Β· 2g P Β· 14.8g C Β· 9.8g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:06:16.616Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Log chicken and yogurt." (difficulty β)0.9s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Decisive guess with stated assumption + invited correction (good-guessing).
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.9s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Chicken breast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount β tell me if that is not right.
3 Β· App data rows written created food_log_entry: Chicken breast Γ1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F created food_log_entry: Plain Greek yogurt Γ1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F created food_log_entry: Chicken breast Γ1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F created food_log_entry: Plain Greek yogurt Γ1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:06:28.650Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I ate a sandwich, some soup, a cookie, milk, and fruit." (difficulty β)2.8s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.8s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve a sandwich, a cookie, and fruit before I log this meal. What should I use for a sandwich, a cookie, and fruit?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:06:42.517Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had a big salad with chicken." (difficulty β)2.5s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.5s
1 Β· TTS saidno speech captured
2 Β· Card shown What exact food and amount should I use for Chicken breast? I did not log it yet because I could not safely finish that food log.
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:06:56.116Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Breakfast: two eggs, toast, bacon, coffee, oatmeal, orange juice, and pancakes." (difficulty β)3.2s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.2s
1 Β· TTS saidno speech captured
2 Β· Card shown I did not make any app changes. What would you like me to do with that?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:07:10.451Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
FAILshould log the entry"Buffet plate: chicken, rice, salad, soup, bread, pasta, fish, vegetables, dessert, and coffee." (difficulty β)1.8s
Verdict Expected LOG β should log the entry. FAIL: OVER-ASK β asked instead of logging (no saved row).
Why verdict OVER-ASK β asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.8s
1 Β· TTS saidno speech captured
2 Β· Card shown I did not make any app changes. What would you like me to do with that?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:07:23.367Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had a ham and swiss on rye and some chips." (difficulty β)2.1s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.1s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve rye before I log this meal. What should I use for rye?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:07:36.534Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Log curry and rice." (difficulty β)2.0s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.0s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve curry before I log this meal. What should I use for curry?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:07:49.632Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"Morning: toast, cereal, yogurt, a protein shake, two eggs, one banana, and coffee." (difficulty β)2.1s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.1s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve cereal, yogurt, and protein shake before I log this meal. What should I use for cereal, yogurt, and protein shake?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:08:02.852Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
UNVERIFIEDunclassifiedshould log the entry"Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana." (difficulty β)3.7s
Verdict Expected LOG β should log the entry. UNVERIFIED: UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
Why verdict UNVERIFIABLE β no saved row captured (untagged / lossy) β resolve via the notes.utterance_trace_id join. App claim not credited; resolve via the notes.utterance_trace_id join.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.7s
1 Β· TTS saidno speech captured
2 Β· Card shown Logging your Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana β still looking up the nutrition details; they'll appear in your log shortly.
3 Β· App data rows written pending_created food_log_entry: Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana (pending)
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:08:24.632Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had pizza, cereal, yogurt, hummus, and one Coca-Cola Zero." (difficulty β)1.9s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.9s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:08:30.665Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould ask a clarifying question"I had a smoothie for breakfast, two eggs, and toast." (difficulty β)2.7s
Verdict Expected CLARIFY β should ask a clarifying question. PASS: Asked a clarification as expected.
Why verdict Asked a clarification as expected.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.7s
1 Β· TTS saidno speech captured
2 Β· Card shown I need to resolve a smoothie before I log this meal. What should I use for a smoothie?
3 Β· App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:08:44.501Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup methodNo lookup method was captured for this path.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould log the entry"Snack was one cup red seedless grapes, one ounce manchego cheese, and one Wasa multigrain crispbread." (difficulty β)1.9s
Verdict Expected LOG β should log the entry. PASS: Logged (write-truth verified): Manchego cheese, Multigrain Crispbread, Grapes β card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.9s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Grapes, one ounce manchego cheese, and Multigrain Crispbread β DIVINA. Assumed 1 oz as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Grapes Γ1 (one cup (151 g)) 104 cal Β· 1.1g P Β· 27.3g C Β· 0.3g F created food_log_entry: Manchego cheese Γ1 (1 oz) 113 cal Β· 7.1g P Β· 0.1g C Β· 9.4g F created food_log_entry: Multigrain Crispbread Γ1 (1 PIECE) 120 cal Β· 4g P Β· 13g C Β· 5g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:08:57.518Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)
PASSshould log the entry"Lunch: one Sweetgreen harvest bowl, one cup miso soup, one sheet nori, one tablespoon sesame seeds, and one cup jasmine green tea." (difficulty β)1.0s
Verdict Expected LOG β should log the entry. PASS: Logged (write-truth verified): Nori, Miso soup, Sweetgreen Harvest Bowl, Green tea, Sesame seeds β card not captured.
Why verdict Logged (write-truth verified): Nori, Miso soup, Sweetgreen Harvest Bowl, Green tea, Sesame seeds β card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS saidno speech captured
2 Β· Card shown Logged Sweetgreen Harvest Bowl, one cup miso soup, one sheet nori, one tablespoon sesame seeds, and one cup jasmine green tea. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Sweetgreen Harvest Bowl Γ1 (bowl) 540 cal Β· 28g P Β· 48g C Β· 24g F created food_log_entry: Miso soup Γ1 (1 cup) 49 cal Β· 3.7g P Β· 6.6g C Β· 1.5g F created food_log_entry: Nori Γ1 (1 sheet nori) 8 cal Β· 1.2g P Β· 1.1g C Β· 0.1g F created food_log_entry: Sesame seeds Γ1 (1 tbsp) 52 cal Β· 1.6g P Β· 2.1g C Β· 4.5g F created food_log_entry: Green tea Γ1 (1 cup) 2 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI didno screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T12:09:09.683Z
6 Β· Why this foodNo food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from a manually curated nutrition entry.
8 Β· Candidates β deciderno candidate list captured (this path did not run DB resolution β e.g. provisional branded log)