compound-dishes-A · preserved partial run · iOS sim
This is the preserved partial result for this run. Its original utterance evidence is shown below.
Rows
13
Pass
11 (85%)
Fail
2 (15%)
Unverified
0 (0%)
Pass rate
85%
Avg difficulty
—
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace — needs a human look (that's you 👍/👎-ing it). "Pass rate" = pass ÷ (pass + fail) — it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows).
Why the fails happened — comprehension vs execution vs cosmetic
Comprehension — picked the wrong action/target (the hard problem)
2 (100%)
resolution — 2 (100% of comprehension)
"Ate an apple, a banana, and grapes." — Compound dish structure mismatch vs expected hierarchy.
"Log a smoothie with one banana, one cup whole milk, and two tablespoons peanut butter." — Compound dish structure mismatch vs expected hierarchy.
Of 2 fails: if most are comprehension, that's the hard problem (the app didn't figure out the right action/target); execution/cosmetic means the decision was right but something downstream broke. Click a bar for the sub-split + example utterances.
Handled correctly? — by expected action
Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed to
N
Correct
Wrong
Unverified
▸ LOG — log the entry
13
11 (85%)
2 (15%)
0 (0%)
2 handled wrong — click one to jump to its full detail below
"Ate an apple, a banana, and grapes."
OVER-ASK — asked instead of logging (no saved row).
"Log a smoothie with one banana, one cup whole milk, and two tablespoons peanut butter."
WRITE-TRUTH FAIL — COMPOUND: component name missing — expected like "milk", got: Peanut butter, Whole milk, Banana
Total
13
11 (85%)
2
0
Accuracy by difficulty
Pending A1's per-utterance difficulty score (requested 2026-07-05) — this bar chart lights up once that lands.
Clarification follow-ups — scored separately
Second turn: app asked, we replied — did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.
Cosmetic only
Not yet classified — pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.
System / infra
Not yet classified — pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.
Latency
Avg (time to ready)
0.8s
p90
1.1s
Max
2.8s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers — click any dot to jump to its detail.
0s
1s
2s
5s
3s
Response path — quick (single response) vs async (an ack like "Working on it…" before the real answer).
Quick response
13
Slowest 8 utterances (click to jump to detail):
"I had a burrito with one tortilla, four ounces steak, half cup rice, and two tablespoons guacamole."2.8s
"I had a turkey sandwich"1.3s
"I had a salad with one cup spinach, one tomato, and two tablespoons balsamic vinaigrette."1.1s
"I had a sandwich with two slices whole wheat bread, three ounces turkey, and one slice cheddar cheese."0.9s
"Ate an apple, a banana, and grapes."0.5s
"Log a protein shake with one banana, two tablespoons peanut butter, and half cup oats."0.5s
"I had a bowl with one cup rice, three ounces chicken, and one cup broccoli."0.5s
"Log a smoothie with one banana, one cup whole milk, and two tablespoons peanut butter."0.5s
Filter — controls the list below
Pass / Fail / Unverified
PASS 11FAIL 2UNVERIFIED 0
Module (intended for)
Other (13)
Utterance sub-type (within module)
13 shown — 11 pass, 2 fail, 0 unverified
Per-utterance detail
PASSshould log the entry"I had an omelet with three eggs, one ounce of cheddar cheese and a tomato" (difficulty —)0.5s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Egg, Cheddar cheese, Omelet, Tomato — card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 · TTS said Logged Omelet. Includes three eggs, one ounce of cheddar cheese, and a tomato.
2 · Card shown Logged Omelet. Includes three eggs, one ounce of cheddar cheese, and a tomato.
3 · App data rows written created food_log_entry: Omelet ×1 (serving) 351 cal · 26.5g P · 6.8g C · 23.9g F created food_log_entry: Egg ×1 (3 eggs) 215 cal · 18.9g P · 1g C · 14.3g F created food_log_entry: Cheddar cheese ×1 (1 oz) 114 cal · 6.5g P · 1g C · 9.4g F created food_log_entry: Tomato ×1 (1 tomato) 22 cal · 1.1g P · 4.8g C · 0.2g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:16:55.956Z
6 · Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 0.0 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had an omelet" (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Omelet — card not captured.
Why verdict Logged (write-truth verified): Omelet — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS said Got it, I logged omelet.
2 · Card shown I logged omelet — let me know if that's not right.
3 · App data rows written created food_log_entry: Omelet ×1 (serving) 351 cal · 26.5g P · 6.8g C · 23.9g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:17:07.296Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
FAILshould log the entry"Ate an apple, a banana, and grapes." (difficulty —)0.5s
Verdict Expected LOG — should log the entry. FAIL: OVER-ASK — asked instead of logging (no saved row).
Why verdict OVER-ASK — asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 · TTS said How much should I log for Grapes? I did not log them yet because the amounts were not clear.
2 · Card shown How much should I log for Grapes? I did not log them yet because the amounts were not clear.
3 · App data rows writtenNo food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:17:18.929Z
6 · Why this food
The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
Matched the common-food list to "Banana" at 90% confidence.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 0.1 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had a sandwich with two slices whole wheat bread, three ounces turkey, and one slice cheddar cheese." (difficulty —)0.9s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Whole wheat bread, Cheddar cheese, Sandwich, Turkey breast — card not captured.
Why verdict Logged (write-truth verified): Whole wheat bread, Cheddar cheese, Sandwich, Turkey breast — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.9s
1 · TTS said Logged Sandwich. Includes two slices whole wheat bread, three ounces turkey, and one slice cheddar cheese.
2 · Card shown Logged Sandwich. Includes two slices whole wheat bread, three ounces turkey, and one slice cheddar cheese.
3 · App data rows written created food_log_entry: Sandwich ×1 (serving) 334 cal · 36.6g P · 23.7g C · 10.5g F created food_log_entry: Whole wheat bread ×1 (2 slices) 138 cal · 7.3g P · 23g C · 2.4g F created food_log_entry: Turkey breast ×1 (3 oz) 115 cal · 24.7g P · 0g C · 1.4g F created food_log_entry: Cheddar cheese ×1 (1 slice) 81 cal · 4.6g P · 0.7g C · 6.7g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:17:30.941Z
6 · Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 0.2 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had a salad with one cup spinach, one tomato, and two tablespoons balsamic vinaigrette." (difficulty —)1.1s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Balsamic vinaigrette, Spinach, Salad, Tomato — card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.1s
1 · TTS said Logged Salad. Includes one cup spinach, one tomato, and two tablespoons balsamic vinaigrette.
2 · Card shown Logged Salad. Includes one cup spinach, one tomato, and two tablespoons balsamic vinaigrette.
3 · App data rows written created food_log_entry: Salad ×1 (serving) 101 cal · 2.1g P · 11.9g C · 5.7g F created food_log_entry: Spinach ×1 (1 cup) 7 cal · 0.9g P · 1.1g C · 0.1g F created food_log_entry: Tomato ×1 (1 tomato) 22 cal · 1.1g P · 4.8g C · 0.2g F created food_log_entry: Balsamic vinaigrette ×1 (2 tbsp) 72 cal · 0.1g P · 6g C · 5.4g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:17:43.307Z
6 · Why this food
The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
Matched the common-food list to "Balsamic vinaigrette" at 90% confidence.
Matched your saved foods and aliases to "a tomato" at 60% confidence.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 0.3 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
FAILshould log the entry"Log a smoothie with one banana, one cup whole milk, and two tablespoons peanut butter." (difficulty —)0.5s
Verdict Expected LOG — should log the entry. FAIL: WRITE-TRUTH FAIL — COMPOUND: component name missing — expected like "milk", got: Peanut butter, Whole milk, Banana
Why verdict WRITE-TRUTH FAIL — COMPOUND: component name missing — expected like "milk", got: Peanut butter, Whole milk, Banana
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 · TTS said Logged Smoothie. Includes one banana, one cup whole milk, and two tablespoons peanut butter.
2 · Card shown Logged Smoothie. Includes one banana, one cup whole milk, and two tablespoons peanut butter.
3 · App data rows written created food_log_entry: Smoothie ×1 (serving) 430 cal · 16.6g P · 44.7g C · 23.6g F created food_log_entry: Banana ×1 (1 banana) 105 cal · 1.3g P · 27.1g C · 0.4g F created food_log_entry: Whole milk ×1 (1 cup) 149 cal · 7.8g P · 11.7g C · 8.1g F created food_log_entry: Peanut butter ×1 (2 tbsp) 176 cal · 7.5g P · 5.9g C · 15.1g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:17:54.797Z
6 · Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 0.3 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had a stir fry with four ounces chicken, one cup broccoli, and one cup rice." (difficulty —)0.4s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Broccoli, Stir Fry, Cooked white rice, Chicken breast — card not captured.
Why verdict Logged (write-truth verified): Broccoli, Stir Fry, Cooked white rice, Chicken breast — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 · TTS said Logged Stir Fry. Includes four ounces chicken, one cup broccoli, and one cup rice.
2 · Card shown Logged Stir Fry. Includes four ounces chicken, one cup broccoli, and one cup rice.
3 · App data rows written created food_log_entry: Stir Fry ×1 (serving) 424 cal · 41.7g P · 50.8g C · 5g F created food_log_entry: Chicken breast ×1 (4 oz) 187 cal · 35.2g P · 0g C · 4.1g F created food_log_entry: Broccoli ×1 (1 cup) 32 cal · 2.2g P · 6.6g C · 0.4g F created food_log_entry: Cooked white rice ×1 (1 cup) 205 cal · 4.3g P · 44.2g C · 0.5g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:18:12.404Z
6 · Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 0.4 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had a burrito with one tortilla, four ounces steak, half cup rice, and two tablespoons guacamole." (difficulty —)2.8s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Guacamole, Cooked white rice, Pork, Steak, Lean Only Eaten, Tortilla, Burrito — card not captured.
Why verdict Logged (write-truth verified): Guacamole, Cooked white rice, Pork, Steak, Lean Only Eaten, Tortilla, Burrito — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 2.8s
1 · TTS said Logged Burrito. Includes one tortilla, Pork, Steak, Lean Only Eaten, half cup rice, and Guacamole.
2 · Card shown Logged Burrito. Includes one tortilla, Pork, Steak, Lean Only Eaten, half cup rice, and Guacamole.
3 · App data rows written created food_log_entry: Burrito ×1 (serving) 554 cal · 37.2g P · 47.6g C · 22.8g F created food_log_entry: Tortilla ×1 (1 tortilla) 140 cal · 3.8g P · 22.5g C · 3.6g F created food_log_entry: Pork, Steak, Lean Only Eaten ×1 (four ounces (113.4 g)) 261 cal · 30.3g P · 0g C · 14.5g F created food_log_entry: Cooked white rice ×1 (0.5 cup) 103 cal · 2.1g P · 22.1g C · 0.2g F created food_log_entry: Guacamole ×1 (tablespoon) 50 cal · 1g P · 3g C · 4.5g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:18:26.391Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Log a protein shake with one banana, two tablespoons peanut butter, and half cup oats." (difficulty —)0.5s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Banana, Protein Shake, Dry oats, Peanut butter — card not captured.
Why verdict Logged (write-truth verified): Banana, Protein Shake, Dry oats, Peanut butter — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 · TTS said Logged Protein Shake. Includes one banana, two tablespoons peanut butter, and half cup oats.
2 · Card shown Logged Protein Shake. Includes one banana, two tablespoons peanut butter, and half cup oats.
3 · App data rows written created food_log_entry: Protein Shake ×1 (serving) 437 cal · 15.6g P · 59.5g C · 18.3g F created food_log_entry: Banana ×1 (1 banana) 105 cal · 1.3g P · 27.1g C · 0.4g F created food_log_entry: Peanut butter ×1 (2 tbsp) 176 cal · 7.5g P · 5.9g C · 15.1g F created food_log_entry: Dry oats ×1 (0.5 cup) 156 cal · 6.8g P · 26.5g C · 2.8g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:18:38.051Z
6 · Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 1.4 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had a bowl with one cup rice, three ounces chicken, and one cup broccoli." (difficulty —)0.5s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Cooked white rice, Broccoli, Chicken breast, Bowl — card not captured.
Why verdict Logged (write-truth verified): Cooked white rice, Broccoli, Chicken breast, Bowl — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 · TTS said Logged Bowl. Includes one cup rice, three ounces chicken, and one cup broccoli.
2 · Card shown Logged Bowl. Includes one cup rice, three ounces chicken, and one cup broccoli.
3 · App data rows written created food_log_entry: Bowl ×1 (serving) 377 cal · 32.9g P · 50.8g C · 4g F created food_log_entry: Cooked white rice ×1 (1 cup) 205 cal · 4.3g P · 44.2g C · 0.5g F created food_log_entry: Chicken breast ×1 (3 oz) 140 cal · 26.4g P · 0g C · 3.1g F created food_log_entry: Broccoli ×1 (1 cup) 32 cal · 2.2g P · 6.6g C · 0.4g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:18:49.892Z
6 · Why this food
The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
Matched your saved foods and aliases to "one cup broccoli" at 90% confidence.
Matched the common-food list to "Broccoli" at 90% confidence.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 0.6 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had a turkey sandwich" (difficulty —)1.3s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Sandwich — card not captured.
Why verdict Logged (write-truth verified): Sandwich — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 1.3s
1 · TTS said Logged sandwich with an estimated macro range.
2 · Card shown Logged sandwich — the macros are an estimate since this food varies by brand.
3 · App data rows written created food_log_entry: Sandwich ×1 (serving) 334 cal · 36.6g P · 23.7g C · 10.5g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:19:02.123Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Dinner was a burrito" (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Burrito — card not captured.
Why verdict Logged (write-truth verified): Burrito — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS said Logged burrito with an estimated macro range.
2 · Card shown Logged burrito — the macros are an estimate since this food varies by brand.
3 · App data rows written created food_log_entry: Burrito ×1 (serving) 554 cal · 37.2g P · 47.6g C · 22.8g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:19:13.429Z
6 · Why this foodNo food-decision explanation was captured for this path.
7 · Lookup methodNo lookup method was captured for this path.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"Breakfast was an omelet with two eggs, one ounce cheddar, and one teaspoon olive oil." (difficulty —)0.5s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Cheddar cheese, Omelet, Olive oil, Egg — card not captured.
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 · TTS said Logged Omelet. Includes two eggs, one ounce cheddar, and one teaspoon olive oil.
2 · Card shown Logged Omelet. Includes two eggs, one ounce cheddar, and one teaspoon olive oil.
3 · App data rows written created food_log_entry: Omelet ×1 (serving) 297 cal · 19.1g P · 1.7g C · 23.4g F created food_log_entry: Egg ×1 (2 eggs) 143 cal · 12.6g P · 0.7g C · 9.5g F created food_log_entry: Cheddar cheese ×1 (1 oz) 114 cal · 6.5g P · 1g C · 9.4g F created food_log_entry: Olive oil ×1 (1 tsp) 40 cal · 0g P · 0g C · 4.5g F
4 · UI didno screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-02T21:19:25.051Z
6 · Why this food The lookup compared against planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
7 · Lookup method
The report captured the app food lookup that ran before or alongside the final decision.
It searched planned foods, your saved foods and aliases, your past food logs, the shared nutrition database, the common-food list, the exercise list, and the workout equipment list.
The lookup evidence took 0.8 seconds to collect.
This lookup is supporting evidence in the report; by itself it does not prove that food was saved.
8 · Candidates → deciderno candidate list captured (this path did not run DB resolution — e.g. provisional branded log)