πŸ“Š Status Dashboard ↑ all runs

multifood-corpus-a-7c039dd0-c16-20…

multifood-corpus-a Β· 57 minutes ago Β· iOS sim
Rows
60
Pass
6 (10%)
Fail
54 (90%)
Unverified
0 (0%)
Pass rate
10%
Avg difficulty
β€”
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace β€” needs a human look (that's you πŸ‘/πŸ‘Ž-ing it). "Pass rate" = pass Γ· (pass + fail) β€” it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows).

Why the fails happened β€” comprehension vs execution vs cosmetic

Comprehension β€” picked the wrong action/target (the hard problem)
53 (98%)
Execution β€” right decision, output broke (plumbing)
1 (2%)
Of 54 fails: if most are comprehension, that's the hard problem (the app didn't figure out the right action/target); execution/cosmetic means the decision was right but something downstream broke. Click a bar for the sub-split + example utterances.

Handled correctly? β€” by expected action

Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed toNCorrectWrongUnverified
β–Έ LOG β€” log the food 606 (10%) 54 (90%) 0 (0%)
Total606 (10%)540

Accuracy by difficulty

Pending A1's per-utterance difficulty score (requested 2026-07-05) β€” this bar chart lights up once that lands.

Clarification follow-ups β€” scored separately

Second turn: app asked, we replied β€” did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.

Cosmetic only

Not yet classified β€” pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.

System / infra

Not yet classified β€” pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.

Latency

Avg (time to ready)
2.3s
p90
4.9s
Max
9.3s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers β€” click any dot to jump to its detail.
0s
1s
2s
5s
10s
Response path β€” quick (single response) vs async (an ack like "Working on it…" before the real answer).
Quick response
60
Slowest 8 utterances (click to jump to detail):
"Track one cup overnight oats with chia, one scoop Vital Proteins collagen peptides, one cup soy milk, one tablespoon maple syrup, one medium blood orange, one ounce hemp hearts, and one cup plain kefir."9.3s
"Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana."9.3s
"I had one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios."7.3s
"Dinner was steak, pasta, wine, bread, and dessert."5.8s
"Breakfast: two eggs, toast, bacon, coffee, oatmeal, orange juice, and pancakes."5.8s
"Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice."5.2s
"I ate one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream."4.9s
"Track toast, one cup Daisy 2 percent cottage cheese, one boiled egg, one cup cantaloupe, one ounce cashews, one Diet Dr Pepper, and one square Lindt 70 percent chocolate."4.9s

Filter β€” controls the list below

Pass / Fail / Unverified
PASS 6 FAIL 54 UNVERIFIED 0
Module (intended for)
food (60)
Utterance sub-type (within module)
60 shown β€” 6 pass, 54 fail, 0 unverified

Per-utterance detail

FAILshould log the food"Log a medium Fuji apple and one KIND peanut butter dark chocolate bar." (difficulty β€”)1.1s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” WRONG/MISSING item "KIND peanut butter dark chocolate bar" β€” no saved row with matching identity (rows: Apple, Chocolate Candy, Peanut Butter Filled)
Why verdict WRITE-TRUTH FAIL β€” WRONG/MISSING item "KIND peanut butter dark chocolate bar" β€” no saved row with matching identity (rows: Apple, Chocolate Candy, Peanut Butter Filled)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.1s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Apple and Chocolate Candy, Peanut Butter Filled. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
created food_log_entry: Chocolate Candy, Peanut Butter Filled Γ—1 (1 large/king size) 438 cal Β· 7g P Β· 57.7g C Β· 22.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:46:11.444Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had a Chipotle chicken burrito bowl and a bottle of Topo Chico." (difficulty β€”)3.1s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.1s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve a Chipotle chicken burrito bowl before I log this meal. What should I use for a Chipotle chicken burrito bowl?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:46:25.731Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Breakfast was two scrambled eggs, one slice of sourdough toast, and eight ounces of cold brew." (difficulty β€”)0.6s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” IMPLAUSIBLE: 265 kcal is implausibly HIGH for 1 count of Toast (typical 80-130 kcal [per-serving])
Why verdict WRITE-TRUTH FAIL β€” IMPLAUSIBLE: 265 kcal is implausibly HIGH for 1 count of Toast (typical 80-130 kcal [per-serving])
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS said no speech captured
2 Β· Card shown Logged two scrambled eggs, Toast, and cold brew. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Egg Γ—1 (2 eggs) 143 cal Β· 12.6g P Β· 0.7g C Β· 9.5g F
created food_log_entry: Toast Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
created food_log_entry: cold brew Γ—1 (serving) 150 cal Β· 10g P Β· 15g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:46:37.404Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Track one Oikos triple zero vanilla yogurt, a medium navel orange, and twelve almonds." (difficulty β€”)0.7s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” WRONG/MISSING item "Oikos triple zero vanilla yogurt" β€” no saved row with matching identity (rows: Orange, Plain Greek yogurt, Almonds)
Why verdict WRITE-TRUTH FAIL β€” WRONG/MISSING item "Oikos triple zero vanilla yogurt" β€” no saved row with matching identity (rows: Orange, Plain Greek yogurt, Almonds)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.7s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Plain Greek yogurt, Orange, and twelve almonds. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Plain Greek yogurt Γ—1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
created food_log_entry: Orange Γ—1 (184 g) 86 cal Β· 1.7g P Β· 21.7g C Β· 0.2g F
created food_log_entry: Almonds Γ—1 (12 almonds) 83 cal Β· 3.1g P Β· 3.1g C Β· 7.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:46:49.218Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the food"Lunch was six ounces grilled Atlantic salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge." (difficulty β€”)0.5s
Verdict Expected LOG β€” should log the food. PASS: Logged (write-truth verified): Broccoli, Lemon, Cooked brown rice, Salmon, Olive oil β€” card not captured.
Why verdict Logged (write-truth verified): Broccoli, Lemon, Cooked brown rice, Salmon, Olive oil β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Salmon, one cup steamed broccoli, half cup brown rice, one tablespoon olive oil, and a lemon wedge. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Salmon Γ—1 (six ounces (170.1 g)) 354 cal Β· 34g P Β· 0g C Β· 22.1g F
created food_log_entry: Broccoli Γ—1 (1 cup) 55 cal Β· 3.8g P Β· 11.3g C Β· 0.6g F
created food_log_entry: Cooked brown rice Γ—1 (0.5 cup) 109 cal Β· 2.2g P Β· 22.9g C Β· 0.8g F
created food_log_entry: Olive oil Γ—1 (1 tbsp) 120 cal Β· 0g P Β· 0g C Β· 13.6g F
created food_log_entry: Lemon Γ—1 (1 lemon wedge) 17 cal Β· 0.6g P Β· 5.4g C Β· 0.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:47:00.900Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I ate a Panera Mediterranean bowl, one Perfect Bar almond butter, a cup of bone broth, two celery stalks, and one Medjool date." (difficulty β€”)1.0s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving])
Why verdict WRITE-TRUTH FAIL β€” IMPLAUSIBLE: 590 kcal is implausibly HIGH for 1 count of Panera Bread Mediterranean Bowl (typical 80-130 kcal [per-serving]); IMPLAUSIBLE: 320 kcal is implausibly HIGH for 1 count of Almond Butter (typical 170-210 kcal [per-serving])
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Panera Bread Mediterranean Bowl, Almond Butter β€” PERFECT BAR, a cup of bone broth, two celery stalks, and one medjool date. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Panera Bread Mediterranean Bowl Γ—1 (1 Bowl) 590 cal Β· 16g P Β· 66g C Β· 31g F
created food_log_entry: Almond Butter Γ—1 (1 BAR) 320 cal Β· 13g P Β· 25g C Β· 19g F
created food_log_entry: Bone broth Γ—1 (1 cup) 36 cal Β· 7.2g P Β· 1.2g C Β· 0.7g F
created food_log_entry: Celery Γ—1 (2 celery stalks) 13 cal Β· 0.6g P Β· 2.4g C Β· 0.2g F
created food_log_entry: Dates Γ—1 (1 medjool date) 66 cal Β· 0.4g P Β· 18g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:47:12.966Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from restaurant menu.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the food"Dinner: eight ounces lean ground turkey, one cup roasted Brussels sprouts, half avocado, one Thomas whole wheat English muffin, one Laughing Cow wedge, ten baby carrots, and sixteen ounces sparkling water." (difficulty β€”)1.2s
Verdict Expected LOG β€” should log the food. PASS: Logged (write-truth verified): Baby carrots, Avocado, Whole wheat English muffin, Brussels sprouts, Water, Lean ground turkey, Water, Laughing Cow β€” card not captured.
Why verdict Logged (write-truth verified): Baby carrots, Avocado, Whole wheat English muffin, Brussels sprouts, Water, Lean ground turkey, Water, Laughing Cow β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.2s
1 Β· TTS said no speech captured
2 Β· Card shown Logged eight ounces lean ground turkey, one cup roasted brussels sprouts, half avocado, Whole wheat English muffin, one laughing cow wedge, ten baby carrots, and Water. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Lean ground turkey Γ—1 (8 oz) 386 cal Β· 61.2g P Β· 0g C Β· 15.9g F
created food_log_entry: Brussels sprouts Γ—1 (1 cup) 70 cal Β· 5.4g P Β· 14.4g C Β· 0.5g F
created food_log_entry: Avocado Γ—1 (half avocado) 120 cal Β· 1.5g P Β· 6.4g C Β· 11g F
created food_log_entry: Whole wheat English muffin Γ—1 (57 g) 128 cal Β· 5.6g P Β· 24.8g C Β· 1.3g F
created food_log_entry: Laughing Cow Γ—1 (1 laughing cow wedge) 35 cal Β· 2g P Β· 1g C Β· 2.5g F
created food_log_entry: Baby carrots Γ—1 (10 baby carrots) 35 cal Β· 0.6g P Β· 8.2g C Β· 0.1g F
created food_log_entry: Water Γ—1 (sixteen ounces (453.6 g)) 0 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:47:25.415Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the food"Log a Starbucks tall oat milk latte, one Kodiak Cakes blueberry waffle, two strips turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and one Justin maple almond butter squeeze pack." (difficulty β€”)1.2s
Verdict Expected LOG β€” should log the food. PASS: Logged (write-truth verified): Oat milk latte, Turkey bacon, Egg, Mozzarella, Blueberry waffle, Almond butter, Blackberries β€” card not captured.
Why verdict Logged (write-truth verified): Oat milk latte, Turkey bacon, Egg, Mozzarella, Blueberry waffle, Almond butter, Blackberries β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.2s
1 Β· TTS said no speech captured
2 Β· Card shown Logged a starbucks tall oat milk latte, one kodiak cakes blueberry waffle, Turkey bacon, one cup blackberries, one hard boiled egg, one ounce part skim mozzarella, and Almond butter. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Oat milk latte Γ—1 (1 starbucks tall oat milk latte) 135 cal Β· 4.3g P Β· 18.5g C Β· 5.3g F
created food_log_entry: Blueberry waffle Γ—1 (1 kodiak cakes blueberry waffle) 180 cal Β· 7g P Β· 24.5g C Β· 4.9g F
created food_log_entry: Turkey bacon Γ—2 (14 g) 60 cal Β· 8.2g P Β· 1g C Β· 3g F
created food_log_entry: Blackberries Γ—1 (1 cup) 62 cal Β· 2g P Β· 14.7g C Β· 0.7g F
created food_log_entry: Egg Γ—1 (1 egg) 72 cal Β· 6.3g P Β· 0.4g C Β· 4.8g F
created food_log_entry: Mozzarella Γ—1 (1 oz) 72 cal Β· 6.9g P Β· 0.8g C Β· 4.5g F
created food_log_entry: Almond butter Γ—1 (32 g) 196 cal Β· 6.7g P Β· 6.1g C Β· 17.9g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:47:37.765Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the food"Snack spread: one Blue Diamond smokehouse almonds pack, one Babybel light, one cup edamame, one Pink Lady apple, one cup coconut water, one Good Culture strawberry cottage cheese cup, one rice cake, one tablespoon chia seeds, five walnuts, and one Clif Kid Zbar." (difficulty β€”)1.9s
Verdict Expected LOG β€” should log the food. PASS: Logged (write-truth verified): Walnuts, Clif Kid Zbar, Babybel light, Apple, Chia seeds, Coconut water, Cottage cheese, Plain rice cakes, Edamame, Almonds β€” card not captured.
Why verdict Logged (write-truth verified): Walnuts, Clif Kid Zbar, Babybel light, Apple, Chia seeds, Coconut water, Cottage cheese, Plain rice cakes, Edamame, Almonds β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.9s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Almonds, one babybel light, one cup edamame, Apple, one cup coconut water, Cottage cheese, one rice cake, one tablespoon chia seeds, five walnuts, and one clif kid zbar. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Almonds Γ—1 (28 g) 162 cal Β· 5.9g P Β· 6g C Β· 14g F
created food_log_entry: Babybel light Γ—1 (1 babybel light) 42 cal Β· 6g P Β· 0g C Β· 2.5g F
created food_log_entry: Edamame Γ—1 (1 cup) 188 cal Β· 18.4g P Β· 13.8g C Β· 8.1g F
created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
created food_log_entry: Coconut water Γ—1 (1 cup) 46 cal Β· 1.7g P Β· 8.9g C Β· 0.5g F
created food_log_entry: Cottage cheese Γ—1 (1 cup) 235 cal Β· 26.6g P Β· 8.2g C Β· 10.3g F
created food_log_entry: Plain rice cakes Γ—1 (1 rice cake) 35 cal Β· 0.7g P Β· 7.3g C Β· 0.3g F
created food_log_entry: Chia seeds Γ—1 (1 tbsp) 51 cal Β· 1.7g P Β· 4.4g C Β· 3.2g F
created food_log_entry: Walnuts Γ—1 (5 walnuts) 65 cal Β· 1.5g P Β· 1.4g C Β· 6.5g F
created food_log_entry: Clif Kid Zbar Γ—1 (1 clif kid zbar) 130 cal Β· 3g P Β· 23g C Β· 4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:47:50.815Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Post-workout: one scoop Optimum Nutrition vanilla whey, one frozen banana, one cup frozen mixed berries, one cup unsweetened almond milk, one tablespoon flaxseed, two Medjool dates, one ounce Ghirardelli 72 percent dark chocolate, one Quaker maple brown sugar instant oatmeal packet, one tablespoon honey, and ice." (difficulty β€”)5.2s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 5.2s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one cup frozen mixed berries and ice before I log this meal. What should I use for one cup frozen mixed berries and ice?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:48:07.182Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had toast and a Chobani peach Greek yogurt." (difficulty β€”)0.4s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "toast (bread type)" as "Toast" (corpus said clarify/ask, must not log); WRONG/MISSING item "Chobani peach Greek yogurt" β€” no saved row with matching identity (rows: Toast, Plain Greek yogurt)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "toast (bread type)" as "Toast" (corpus said clarify/ask, must not log); WRONG/MISSING item "Chobani peach Greek yogurt" β€” no saved row with matching identity (rows: Toast, Plain Greek yogurt)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Toast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Toast Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
created food_log_entry: Plain Greek yogurt Γ—1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:48:18.781Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Log cottage cheese and one large Honeycrisp apple." (difficulty β€”)0.6s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "cottage cheese (fat or brand)" as "Cottage cheese" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "cottage cheese (fat or brand)" as "Cottage cheese" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Cottage cheese and Apple. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Cottage cheese Γ—1 (100 g) 98 cal Β· 11.1g P Β· 3.4g C Β· 4.3g F
created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:48:30.516Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I ate a turkey sandwich on rye, some chips, and a Diet Coke." (difficulty β€”)3.1s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.1s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve a turkey sandwich on rye before I log this meal. What should I use for a turkey sandwich on rye?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:48:44.865Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Breakfast was two eggs, toast, and black coffee." (difficulty β€”)2.4s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.4s
1 Β· TTS said no speech captured
2 Β· Card shown I did not make any app changes. What would you like me to do with that?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:48:58.365Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had a protein shake, one RXBAR chocolate sea salt bar, and a banana." (difficulty β€”)0.7s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "protein shake (brand or recipe)" as "Protein powder" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "protein shake (brand or recipe)" as "Protein powder" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.7s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Protein powder, Chocolate Sea Salt Bar β€” RXBAR, and a banana. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Protein powder Γ—1 (100 g) 367 cal Β· 83g P Β· 3.3g C Β· 1.7g F
created food_log_entry: Chocolate Sea Salt Bar Γ—1 (1 bar) 200 cal Β· 12g P Β· 23g C Β· 8g F
created food_log_entry: Banana Γ—1 (1 banana) 105 cal Β· 1.3g P Β· 27.1g C Β· 0.4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:49:10.214Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Lunch was a Chipotle steak burrito, granola, one Fuji apple, one cup sparkling water, and a string cheese." (difficulty β€”)2.2s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.2s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve a Chipotle steak burrito before I log this meal. What should I use for a Chipotle steak burrito?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:49:23.582Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had six ounces grilled salmon, peanut butter, one cup steamed asparagus, one Diet Pepsi, and one ounce sharp cheddar." (difficulty β€”)0.6s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "peanut butter (type or brand)" as "Peanut butter" (corpus said clarify/ask, must not log); WRONG/MISSING item "Diet Pepsi" β€” no saved row with matching identity (rows: Asparagus, Cheddar cheese, Salmon, Diet cola, Peanut butter)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "peanut butter (type or brand)" as "Peanut butter" (corpus said clarify/ask, must not log); WRONG/MISSING item "Diet Pepsi" β€” no saved row with matching identity (rows: Asparagus, Cheddar cheese, Salmon, Diet cola, Peanut butter)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS said no speech captured
2 Β· Card shown Logged six ounces grilled salmon, Peanut butter, one cup steamed asparagus, Diet cola, and Cheddar cheese. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Salmon Γ—1 (6 oz) 354 cal Β· 34g P Β· 0g C Β· 22.1g F
created food_log_entry: Peanut butter Γ—1 (32 g) 188 cal Β· 8g P Β· 6.3g C Β· 16.1g F
created food_log_entry: Asparagus Γ—1 (1 cup) 40 cal Β· 4.3g P Β· 7.4g C Β· 0.4g F
created food_log_entry: Diet cola Γ—1 (355 g) 0 cal Β· 0g P Β· 0g C Β· 0g F
created food_log_entry: Cheddar cheese Γ—1 (one ounce (28.3 g)) 114 cal Β· 6.5g P Β· 1g C Β· 9.4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:49:35.365Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Track a Wendy's spicy chicken sandwich, cereal, one cup two percent milk, one clementine, and one ounce provolone." (difficulty β€”)1.3s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:49:47.815Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Dinner was chicken, one cup quinoa, one roasted sweet potato, one tablespoon butter, one Quest chocolate chip cookie dough bar, one cup green tea, and one square Ghirardelli intense dark 86 percent." (difficulty β€”)4.5s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 4.5s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one square Ghirardelli intense dark 86 percent before I log this meal. What should I use for one square Ghirardelli intense dark 86 percent?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:50:03.498Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I ate pizza, one cup Greek salad with feta, one bottle La Croix lime, two eggs, one cup vanilla bean ice cream, one slice whole grain bread, and six ounces Atlantic cod." (difficulty β€”)1.0s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "pizza (type and portion)" as "Pizza" (corpus said clarify/ask, must not log); WRONG/MISSING item "Greek salad with feta" β€” no saved row with matching identity (rows: Cod, Egg, Feta cheese, Vanilla ice cream, Pizza, Lime, Bread); IMPLAUSIBLE: 265 kcal is implausibly HIGH for 1 count of Bread (typical 80-130 kcal [per-serving])
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "pizza (type and portion)" as "Pizza" (corpus said clarify/ask, must not log); WRONG/MISSING item "Greek salad with feta" β€” no saved row with matching identity (rows: Cod, Egg, Feta cheese, Vanilla ice cream, Pizza, Lime, Bread); IMPLAUSIBLE: 265 kcal is implausibly HIGH for 1 count of Bread (typical 80-130 kcal [per-serving])
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Pizza, Feta cheese, Lime, two eggs, Vanilla ice cream, Bread, and Cod. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Pizza Γ—1 (107 g) 285 cal Β· 12.2g P Β· 35.6g C Β· 10.4g F
created food_log_entry: Feta cheese Γ—1 (one cup (150 g)) 396 cal Β· 21.3g P Β· 6.1g C Β· 32g F
created food_log_entry: Lime Γ—1 (67 g) 20 cal Β· 0.5g P Β· 7g C Β· 0.1g F
created food_log_entry: Egg Γ—1 (2 eggs) 143 cal Β· 12.6g P Β· 0.7g C Β· 9.5g F
created food_log_entry: Vanilla ice cream Γ—1 (one cup (240 g)) 497 cal Β· 8.4g P Β· 56.6g C Β· 26.4g F
created food_log_entry: Bread Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
created food_log_entry: Cod Γ—1 (six ounces (170.1 g)) 179 cal Β· 39.1g P Β· 0g C Β· 1.5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:50:15.650Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Log yogurt, one medium Bartlett pear, one ounce pistachios, one In-N-Out grilled cheese, one tablespoon Jif creamy peanut butter, one cup peppermint tea, and one mandarin." (difficulty β€”)3.1s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.1s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one cup peppermint tea and one mandarin before I log this meal. What should I use for one cup peppermint tea and one mandarin?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:50:45.399Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Big breakfast: two eggs, one strip turkey bacon, one medium banana, one cup black coffee, one slice whole wheat bread, six ounces plain nonfat Greek yogurt, and toast." (difficulty β€”)1.4s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "toast (bread type)" as "Toast" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "toast (bread type)" as "Toast" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.4s
1 Β· TTS said no speech captured
2 Β· Card shown Logged two eggs, Turkey bacon, one medium banana, one cup black coffee, Whole wheat bread, six ounces plain nonfat greek yogurt, and Toast. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Egg Γ—1 (2 eggs) 143 cal Β· 12.6g P Β· 0.7g C Β· 9.5g F
created food_log_entry: Turkey bacon Γ—1 (14 g) 30 cal Β· 4.1g P Β· 0.5g C Β· 1.5g F
created food_log_entry: Banana Γ—1 (1 medium banana) 105 cal Β· 1.3g P Β· 27.1g C Β· 0.4g F
created food_log_entry: Coffee Γ—1 (1 cup) 2 cal Β· 0.2g P Β· 0g C Β· 0g F
created food_log_entry: Whole wheat bread Γ—1 (1 slice) 69 cal Β· 3.6g P Β· 11.5g C Β· 1.2g F
created food_log_entry: Nonfat Greek yogurt Γ—1 (6 oz) 100 cal Β· 17.5g P Β· 6.1g C Β· 0.7g F
created food_log_entry: Toast Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:50:57.966Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Meal prep: four ounces chicken breast, one cup jasmine rice, one cup green beans, hummus, one tablespoon olive oil, one ounce feta, one cup plain kefir, one celery stalk, one tablespoon pumpkin seeds, and one cup chamomile tea." (difficulty β€”)0.8s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "hummus (brand or type)" as "Hummus" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "hummus (brand or type)" as "Hummus" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.8s
1 Β· TTS said no speech captured
2 Β· Card shown Logged four ounces chicken breast, one cup jasmine rice, one cup green beans, Hummus, one tablespoon olive oil, one ounce feta, one cup plain kefir, one celery stalk, one tablespoon pumpkin seeds, and one cup chamomile tea. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Chicken breast Γ—1 (4 oz) 187 cal Β· 35.2g P Β· 0g C Β· 4.1g F
created food_log_entry: Cooked jasmine rice Γ—1 (1 cup) 205 cal Β· 3.8g P Β· 44.6g C Β· 0.5g F
created food_log_entry: Green beans Γ—1 (1 cup) 44 cal Β· 2.4g P Β· 9.9g C Β· 0.4g F
created food_log_entry: Hummus Γ—1 (100 g) 166 cal Β· 7.9g P Β· 14.3g C Β· 9.6g F
created food_log_entry: Olive oil Γ—1 (1 tbsp) 120 cal Β· 0g P Β· 0g C Β· 13.6g F
created food_log_entry: Feta cheese Γ—1 (1 oz) 75 cal Β· 4g P Β· 1.2g C Β· 6g F
created food_log_entry: Kefir Γ—1 (1 cup) 100 cal Β· 8.1g P Β· 11g C Β· 2.5g F
created food_log_entry: Celery Γ—1 (1 celery stalk) 6 cal Β· 0.3g P Β· 1.2g C Β· 0.1g F
created food_log_entry: Pumpkin seeds Γ—1 (1 tbsp) 84 cal Β· 4.5g P Β· 1.6g C Β· 7.4g F
created food_log_entry: Chamomile tea Γ—1 (1 cup) 2 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:51:09.982Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Log chicken and one medium Granny Smith apple." (difficulty β€”)0.5s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "chicken (cut and prep)" as "Chicken breast" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "chicken (cut and prep)" as "Chicken breast" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Chicken breast and Apple. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Chicken breast Γ—1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F
created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:51:21.749Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had ice cream and one pint Halo Top sea salt caramel." (difficulty β€”)0.8s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "ice cream (flavor or brand)" as "Vanilla ice cream" (corpus said clarify/ask, must not log); WRONG/MISSING item "Halo Top sea salt caramel pint" β€” no saved row with matching identity (rows: Vanilla ice cream, Sea Salt Caramel Pops)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "ice cream (flavor or brand)" as "Vanilla ice cream" (corpus said clarify/ask, must not log); WRONG/MISSING item "Halo Top sea salt caramel pint" β€” no saved row with matching identity (rows: Vanilla ice cream, Sea Salt Caramel Pops)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.8s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Vanilla ice cream and Sea Salt Caramel Pops β€” HALO TOP. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Vanilla ice cream Γ—1 (100 g) 207 cal Β· 3.5g P Β· 23.6g C Β· 11g F
created food_log_entry: Sea Salt Caramel Pops Γ—1 (1 pop (78g)) 100 cal Β· 6g P Β· 18g C Β· 3g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:51:33.650Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Snack was a protein bar, one cup raspberries, and one mozzarella stick." (difficulty β€”)0.4s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "protein bar (brand)" as "Generic protein bar" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "protein bar (brand)" as "Generic protein bar" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 Β· TTS said no speech captured
2 Β· Card shown Logged a protein bar, one cup raspberries, and one mozzarella stick. Assumed 1 protein bar as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Generic protein bar Γ—1 (1 protein bar) 200 cal Β· 20g P Β· 22g C Β· 7g F
created food_log_entry: Raspberries Γ—1 (1 cup) 64 cal Β· 1.5g P Β· 14.6g C Β· 0.9g F
created food_log_entry: Mozzarella stick Γ—1 (1 mozzarella stick) 91 cal Β· 4.2g P Β· 7g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:51:45.264Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I ate some chips, one Subway turkey six inch on wheat, and a bottle of water." (difficulty β€”)1.0s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "chips (brand or type)" as "Potato chips" (corpus said clarify/ask, must not log); WRONG/MISSING item "Subway turkey six inch on wheat" β€” no saved row with matching identity (rows: Subway Oven Roasted Turkey, Water, Potato chips)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "chips (brand or type)" as "Potato chips" (corpus said clarify/ask, must not log); WRONG/MISSING item "Subway turkey six inch on wheat" β€” no saved row with matching identity (rows: Subway Oven Roasted Turkey, Water, Potato chips)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Potato chips, Subway Oven Roasted Turkey, and Water. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Potato chips Γ—1 (28 g) 150 cal Β· 2g P Β· 14.8g C Β· 9.8g F
created food_log_entry: Subway Oven Roasted Turkey Γ—1 (57 g) 60 cal Β· 11g P Β· 0g C Β· 1g F
created food_log_entry: Water Γ—1 (240 g) 0 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:51:57.416Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Lunch: one Chipotle burrito bowl with chicken black beans and fajita veggies, yogurt, one clementine, and one ounce almonds." (difficulty β€”)2.5s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.5s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve a Chipotle burrito bowl before I log this meal. What should I use for a Chipotle burrito bowl?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:52:11.099Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had a McDonald's Big Mac, cereal, one cup skim milk, one kiwi, and one tablespoon peanut butter on celery." (difficulty β€”)0.9s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.9s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:52:23.066Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Track toast, one cup Daisy 2 percent cottage cheese, one boiled egg, one cup cantaloupe, one ounce cashews, one Diet Dr Pepper, and one square Lindt 70 percent chocolate." (difficulty β€”)4.9s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 4.9s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one square Lindt 70 percent chocolate before I log this meal. What should I use for one square Lindt 70 percent chocolate?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:52:39.115Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Dinner was six ounces pork tenderloin, one baked potato with sour cream, one cup peas, one dinner roll, granola, one cup rooibos tea, and one pluot." (difficulty β€”)4.4s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 4.4s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one dinner roll before I log this meal. What should I use for one dinner roll?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:52:54.650Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had peanut butter and one Fairlife chocolate shake." (difficulty β€”)0.6s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.6s
1 Β· TTS said no speech captured
2 Β· Card shown What exact ingredient should I use for one fairlife chocolate?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:53:06.350Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Log a cup of yogurt, one slice Dave's Killer Bread 21 whole grains, and one cup green tea." (difficulty β€”)1.1s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "yogurt (style or flavor)" as "Plain Greek yogurt" (corpus said clarify/ask, must not log); IMPLAUSIBLE: 265 kcal is implausibly HIGH for 1 count of Bread (typical 80-130 kcal [per-serving])
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "yogurt (style or flavor)" as "Plain Greek yogurt" (corpus said clarify/ask, must not log); IMPLAUSIBLE: 265 kcal is implausibly HIGH for 1 count of Bread (typical 80-130 kcal [per-serving])
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.1s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Plain Greek yogurt, Bread, and one cup green tea. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Plain Greek yogurt Γ—1 (a cup (245 g)) 238 cal Β· 22.1g P Β· 9.6g C Β· 12.3g F
created food_log_entry: Bread Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
created food_log_entry: Green tea Γ—1 (1 cup) 2 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:53:18.599Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Afternoon: one Taco Bell crunchy taco supreme, hummus, one cup pineapple chunks, one tablespoon ranch on carrots, and one Siggi's vanilla skyr." (difficulty β€”)1.4s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "hummus (brand or type)" as "Hummus" (corpus said clarify/ask, must not log); WRONG/MISSING item "ranch on carrots" β€” no saved row with matching identity (rows: Pineapple, Taco Bell Crunchy Taco Supreme, Skyr, Baby carrots, Hummus)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "hummus (brand or type)" as "Hummus" (corpus said clarify/ask, must not log); WRONG/MISSING item "ranch on carrots" β€” no saved row with matching identity (rows: Pineapple, Taco Bell Crunchy Taco Supreme, Skyr, Baby carrots, Hummus)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.4s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Taco Bell Crunchy Taco Supreme, Hummus, one cup pineapple chunks, Baby carrots, and Skyr. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Taco Bell Crunchy Taco Supreme Γ—1 (1 menu item) 190 cal Β· 8g P Β· 15g C Β· 11g F
created food_log_entry: Hummus Γ—1 (100 g) 166 cal Β· 7.9g P Β· 14.3g C Β· 9.6g F
created food_log_entry: Pineapple Γ—1 (1 cup) 83 cal Β· 0.8g P Β· 21.6g C Β· 0.2g F
created food_log_entry: Baby carrots Γ—1 (one tablespoon (9.3 g)) 3 cal Β· 0.1g P Β· 0.8g C Β· 0g F
created food_log_entry: Skyr Γ—1 (170 g) 107 cal Β· 18.7g P Β· 6.8g C Β· 0.3g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:53:31.134Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from restaurant menu.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I ate a Panera broccoli cheddar soup bread bowl, one clementine, a protein shake, one cup cucumber slices, one ounce swiss cheese, one cup peppermint tea, and one madeleine cookie." (difficulty β€”)3.2s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.2s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one madeleine cookie before I log this meal. What should I use for one madeleine cookie?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:53:45.433Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Road trip: one bag Snyder's pretzels, one bottle Gatorade fruit punch, one ounce Jack Link's beef jerky, one Envy apple, one cheese stick, one cup grapes, and toast." (difficulty β€”)2.1s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "toast (bread type)" as "Toast" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "toast (bread type)" as "Toast" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.1s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Pretzels, Fruit Punch Thirst Quencher β€” GATORADE, Beef jerky, Apple, one cheese stick, one cup grapes, and Toast. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Pretzels Γ—1 (28 g) 106 cal Β· 2.5g P Β· 22.1g C Β· 1g F
created food_log_entry: Fruit Punch Thirst Quencher Γ—1 (1 Bottle) 142 cal Β· 0g P Β· 36.1g C Β· 0g F
created food_log_entry: Beef jerky Γ—1 (one ounce (28.3 g)) 116 cal Β· 9.4g P Β· 3.1g C Β· 7.1g F
created food_log_entry: Apple Γ—1 (182 g) 95 cal Β· 0.5g P Β· 25.5g C Β· 0.4g F
created food_log_entry: Mozzarella stick Γ—1 (1 cheese stick) 91 cal Β· 4.2g P Β· 7g C Β· 5g F
created food_log_entry: Grapes Γ—1 (1 cup) 104 cal Β· 1.1g P Β· 27.3g C Β· 0.3g F
created food_log_entry: Toast Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:53:58.866Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I drank one can La Croix pamplemousse and ate one pack SkinnyPop original." (difficulty β€”)3.6s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS said no speech captured
2 Β· Card shown What should I use for one can La Croix pamplemousse? I did not log it yet because I could not match it safely.
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:54:13.650Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had one cup nonfat cottage cheese with pineapple, one slice Ezekiel bread, and twelve pistachios." (difficulty β€”)7.3s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 7.3s
1 Β· TTS said no speech captured
2 Β· Card shown Was it plain, or did you have anything on it like butter or cream cheese? I did not log it yet because add-ons can change the nutrition.
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:54:32.101Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Log one Whataburger honey butter chicken biscuit, one cup tomato soup, three ounces deli turkey, one cup arugula salad, and one tablespoon balsamic vinegar." (difficulty β€”)3.6s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 3.6s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one Whataburger honey butter chicken biscuit before I log this meal. What should I use for one Whataburger honey butter chicken biscuit?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:54:46.834Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I ate one cup pad thai, one vegetable spring roll, one cup mango sticky rice, one bottle San Pellegrino, and one mochi green tea ice cream." (difficulty β€”)4.9s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 4.9s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one bottle San Pellegrino and one mochi green tea ice cream before I log this meal. What should I use for one bottle San Pellegrino and one mochi green tea ice cream?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:55:02.917Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Meal: one Impossible Whopper from Burger King, one medium sweet potato fries, one side garden salad no dressing, one cup unsweetened iced tea, one ounce pepper jack, one cup sauerkraut, and one tablespoon mustard." (difficulty β€”)0.1s
Verdict Expected LOG β€” should log the food. FAIL: Reported could-not-confirm (write-confirmation false negative).
Why verdict Reported could-not-confirm (write-confirmation false negative).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.1s
1 Β· TTS said no speech captured
2 Β· Card shown I could not confirm that the change was saved, so I did not mark it as done.
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:55:29.801Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Track one cup overnight oats with chia, one scoop Vital Proteins collagen peptides, one cup soy milk, one tablespoon maple syrup, one medium blood orange, one ounce hemp hearts, and one cup plain kefir." (difficulty β€”)9.3s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 9.3s
1 Β· TTS said no speech captured
2 Β· Card shown What should I use for one ounce hemp hearts?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:55:50.251Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Full plate: one cup lentil soup, four ounces baked cod, one cup roasted cauliflower, half cup farro, one tablespoon tahini, one cup kale salad, one ounce goat cheese, one tablespoon dried cranberries, one cup Health-Ade ginger lemon kombucha, and one fresh fig." (difficulty β€”)2.0s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.0s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one cup kale salad before I log this meal. What should I use for one cup kale salad?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:56:03.384Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Log curry and rice with a side Caesar salad." (difficulty β€”)1.5s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.5s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve curry before I log this meal. What should I use for curry?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:56:16.066Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had a smoothie, a salad, and soup." (difficulty β€”)2.8s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.8s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve a smoothie and a salad before I log this meal. What should I use for a smoothie and a salad?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:56:30.033Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Dinner was steak, pasta, wine, bread, and dessert." (difficulty β€”)5.8s
Verdict Expected LOG β€” should log the food. FAIL: Reported could-not-confirm (write-confirmation false negative).
Why verdict Reported could-not-confirm (write-confirmation false negative).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 5.8s
1 Β· TTS said no speech captured
2 Β· Card shown What would you like me to do with that?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:56:46.984Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had toast and chips." (difficulty β€”)0.4s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "toast (bread type)" as "Toast" (corpus said clarify/ask, must not log); OVER-ACCEPT β€” logged clarify-item "chips (brand or type)" as "Potato chips" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "toast (bread type)" as "Toast" (corpus said clarify/ask, must not log); OVER-ACCEPT β€” logged clarify-item "chips (brand or type)" as "Potato chips" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Toast and Potato chips. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Toast Γ—1 (100 g) 265 cal Β· 9g P Β· 49g C Β· 3.2g F
created food_log_entry: Potato chips Γ—1 (28 g) 150 cal Β· 2g P Β· 14.8g C Β· 9.8g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:56:58.482Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Log chicken and yogurt." (difficulty β€”)0.9s
Verdict Expected LOG β€” should log the food. FAIL: WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "chicken (cut and prep)" as "Chicken breast" (corpus said clarify/ask, must not log); OVER-ACCEPT β€” logged clarify-item "yogurt (style or flavor)" as "Plain Greek yogurt" (corpus said clarify/ask, must not log)
Why verdict WRITE-TRUTH FAIL β€” OVER-ACCEPT β€” logged clarify-item "chicken (cut and prep)" as "Chicken breast" (corpus said clarify/ask, must not log); OVER-ACCEPT β€” logged clarify-item "yogurt (style or flavor)" as "Plain Greek yogurt" (corpus said clarify/ask, must not log)
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.9s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Chicken breast and Plain Greek yogurt. Assumed catalog default servings where you did not say an amount β€” tell me if that is not right.
3 Β· App data rows written created food_log_entry: Chicken breast Γ—1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F
created food_log_entry: Plain Greek yogurt Γ—1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
created food_log_entry: Chicken breast Γ—1 (100 g) 165 cal Β· 31g P Β· 0g C Β· 3.6g F
created food_log_entry: Plain Greek yogurt Γ—1 (100 g) 97 cal Β· 9g P Β· 3.9g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:57:10.501Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I ate a sandwich, some soup, a cookie, milk, and fruit." (difficulty β€”)2.4s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.4s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve a sandwich, a cookie, and fruit before I log this meal. What should I use for a sandwich, a cookie, and fruit?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:57:24.035Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had a big salad with chicken." (difficulty β€”)2.5s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.5s
1 Β· TTS said no speech captured
2 Β· Card shown What exact food and amount should I use for Chicken breast? I did not log it yet because I could not safely finish that food log.
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:57:37.716Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Breakfast: two eggs, toast, bacon, coffee, oatmeal, orange juice, and pancakes." (difficulty β€”)5.8s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 5.8s
1 Β· TTS said no speech captured
2 Β· Card shown What would you like me to do with that?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:57:54.736Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Buffet plate: chicken, rice, salad, soup, bread, pasta, fish, vegetables, dessert, and coffee." (difficulty β€”)2.3s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.3s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve one buffet plate before I log this meal. What should I use for one buffet plate?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:58:08.166Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had a ham and swiss on rye and some chips." (difficulty β€”)2.4s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.4s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve rye before I log this meal. What should I use for rye?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:58:21.668Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Log curry and rice." (difficulty β€”)1.9s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.9s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve curry before I log this meal. What should I use for curry?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:58:34.785Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Morning: toast, cereal, yogurt, a protein shake, two eggs, one banana, and coffee." (difficulty β€”)2.0s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.0s
1 Β· TTS said no speech captured
2 Β· Card shown I did not make any app changes. What would you like me to do with that?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:58:47.901Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"Ten proteins: chicken, steak, fish, shrimp, lobster, crab, lamb, duck, goose, and one medium banana." (difficulty β€”)9.3s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 9.3s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve fish, lamb, duck, and goose before I log this meal. What should I use for fish, lamb, duck, and goose?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:59:08.451Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had pizza, cereal, yogurt, hummus, and one Coca-Cola Zero." (difficulty β€”)1.0s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve cereal before I log this meal. What should I use for cereal?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:59:20.583Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
FAILshould log the food"I had a smoothie for breakfast, two eggs, and toast." (difficulty β€”)1.3s
Verdict Expected LOG β€” should log the food. FAIL: OVER-ASK β€” asked instead of logging (no saved row).
Why verdict OVER-ASK β€” asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 Β· TTS said no speech captured
2 Β· Card shown I need to resolve a smoothie before I log this meal. What should I use for a smoothie?
3 Β· App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:59:33.016Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method No lookup method was captured for this path.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the food"Snack was one cup red seedless grapes, one ounce manchego cheese, and one Wasa multigrain crispbread." (difficulty β€”)1.6s
Verdict Expected LOG β€” should log the food. PASS: Logged (write-truth verified): Grapes, Manchego cheese, Grapes, Manchego cheese, Multigrain Crispbread β€” card not captured.
Why verdict Logged (write-truth verified): Grapes, Manchego cheese, Grapes, Manchego cheese, Multigrain Crispbread β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.6s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Grapes, one ounce manchego cheese, and Multigrain Crispbread β€” DIVINA. Assumed 1 oz as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Grapes Γ—1 (one cup (151 g)) 104 cal Β· 1.1g P Β· 27.3g C Β· 0.3g F
created food_log_entry: Manchego cheese Γ—1 (1 oz) 113 cal Β· 7.1g P Β· 0.1g C Β· 9.4g F
created food_log_entry: Multigrain Crispbread Γ—1 (1 PIECE) 120 cal Β· 4g P Β· 13g C Β· 5g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:59:45.818Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from the common-food list.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)
PASSshould log the food"Lunch: one Sweetgreen harvest bowl, one cup miso soup, one sheet nori, one tablespoon sesame seeds, and one cup jasmine green tea." (difficulty β€”)1.0s
Verdict Expected LOG β€” should log the food. PASS: Logged (write-truth verified): Sweetgreen Harvest Bowl, Sesame seeds, Miso soup, Nori, Green tea β€” card not captured.
Why verdict Logged (write-truth verified): Sweetgreen Harvest Bowl, Sesame seeds, Miso soup, Nori, Green tea β€” card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.0s
1 Β· TTS said no speech captured
2 Β· Card shown Logged Sweetgreen Harvest Bowl, one cup miso soup, one sheet nori, one tablespoon sesame seeds, and one cup jasmine green tea. Assumed 1 cup as the default serving. Tell me if you want a different amount. Tell me if that is not right.
3 Β· App data rows written created food_log_entry: Sweetgreen Harvest Bowl Γ—1 (bowl) 540 cal Β· 28g P Β· 48g C Β· 24g F
created food_log_entry: Miso soup Γ—1 (1 cup) 49 cal Β· 3.7g P Β· 6.6g C Β· 1.5g F
created food_log_entry: Nori Γ—1 (1 sheet nori) 8 cal Β· 1.2g P Β· 1.1g C Β· 0.1g F
created food_log_entry: Sesame seeds Γ—1 (1 tbsp) 86 cal Β· 2.7g P Β· 3.5g C Β· 7.5g F
created food_log_entry: Green tea Γ—1 (1 cup) 2 cal Β· 0g P Β· 0g C Β· 0g F
4 Β· UI did no screen events captured
5 Β· Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-07-30T19:59:57.918Z
6 Β· Why this food No food-decision explanation was captured for this path.
7 Β· Lookup method The answer used provisional food evidence from a manually curated nutrition entry.
8 Β· Candidates β†’ decider no candidate list captured (this path did not run DB resolution β€” e.g. provisional branded log)