📊 Status Dashboard ↑ all runs

compound-dishes-A-fb954a4d-c43-20260801

compound-dishes-A · preserved partial run · 101s elapsed · iOS sim
This is the preserved partial result for this run. Its original utterance evidence is shown below.
Rows
3
Pass
2 (67%)
Fail
1 (33%)
Unverified
0 (0%)
Pass rate
67%
Avg difficulty
"Unverified" = the grader couldn't confidently call it PASS or FAIL from the trace — needs a human look (that's you 👍/👎-ing it). "Pass rate" = pass ÷ (pass + fail) — it excludes Unverified rows, so it differs slightly from the Pass % above (which is share of ALL rows).

Why the fails happened — comprehension vs execution vs cosmetic

Comprehension — picked the wrong action/target (the hard problem)
1 (100%)
Of 1 fails: if most are comprehension, that's the hard problem (the app didn't figure out the right action/target); execution/cosmetic means the decision was right but something downstream broke. Click a bar for the sub-split + example utterances.

Handled correctly? — by expected action

Click a row to see the ones it got wrong; click a wrong utterance to jump to its full detail below.
Supposed toNCorrectWrongUnverified
▸ LOG — log the entry 32 (67%) 1 (33%) 0 (0%)
Total32 (67%)10

Accuracy by difficulty

Pending A1's per-utterance difficulty score (requested 2026-07-05) — this bar chart lights up once that lands.

Clarification follow-ups — scored separately

Second turn: app asked, we replied — did it resolve correctly?
No CLARIFY_ANSWER (follow-up) rows in this run.

Cosmetic only

Not yet classified — pending A1 adding cosmetic-issue detection (double response, wording variance, rounding) to the grader. Nothing fabricated here.

System / infra

Not yet classified — pending confirmation from A1 whether trace data already carries infra-error signal (TTS/sync/DB-write errors) or needs new detection.

Latency

Avg (time to ready)
0.4s
p90
0.4s
Max
0.5s
Went async
0 (0%)
Each dot is one utterance (green=pass, red=fail, amber=unverified), positioned by how long it took to be ready for review. Bigger dots are the slow outliers — click any dot to jump to its detail.
0s
1s
2s
5s
Response path — quick (single response) vs async (an ack like "Working on it…" before the real answer).
Quick response
3
Slowest 3 utterances (click to jump to detail):
"Ate an apple, a banana, and grapes."0.5s
"I had an omelet with three eggs, one ounce of cheddar cheese and a tomato"0.4s
"I had an omelet"0.2s

Filter — controls the list below

Pass / Fail / Unverified
PASS 2 FAIL 1 UNVERIFIED 0
Module (intended for)
Other (3)
Utterance sub-type (within module)
3 shown — 2 pass, 1 fail, 0 unverified

Per-utterance detail

PASSshould log the entry"I had an omelet with three eggs, one ounce of cheddar cheese and a tomato" (difficulty —)0.4s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Omelet, Egg, Tomato, Cheddar cheese — card not captured.
Why verdict Logged (write-truth verified): Omelet, Egg, Tomato, Cheddar cheese — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.4s
1 · TTS said no speech captured
2 · Card shown Logged Omelet. Includes three eggs, one ounce of cheddar cheese, and a tomato.
3 · App data rows written created food_log_entry: Omelet ×1 (serving) 351 cal · 26.5g P · 6.8g C · 23.9g F
created food_log_entry: Egg ×1 (3 eggs) 215 cal · 18.9g P · 1g C · 14.3g F
created food_log_entry: Cheddar cheese ×1 (1 oz) 114 cal · 6.5g P · 1g C · 9.4g F
created food_log_entry: Tomato ×1 (1 tomato) 22 cal · 1.1g P · 4.8g C · 0.2g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:25:45.815Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from a manually curated nutrition entry.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
PASSshould log the entry"I had an omelet" (difficulty —)0.2s
Verdict Expected LOG — should log the entry. PASS: Logged (write-truth verified): Omelet — card not captured.
Why verdict Logged (write-truth verified): Omelet — card not captured.
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.2s
1 · TTS said no speech captured
2 · Card shown I logged omelet — let me know if that's not right.
3 · App data rows written created food_log_entry: Omelet ×1 (serving) 351 cal · 26.5g P · 6.8g C · 23.9g F
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:25:57.091Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method The answer used provisional food evidence from a manually curated nutrition entry.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)
FAILshould log the entry"Ate an apple, a banana, and grapes." (difficulty —)0.5s
Verdict Expected LOG — should log the entry. FAIL: OVER-ASK — asked instead of logging (no saved row).
Why verdict OVER-ASK — asked instead of logging (no saved row).
My override
Things I think are wrong
Use this when the report passed something that should have failed.
Comprehension
Execution
Cosmetic
Things the report marked wrong but I think are right
Use this when the report failed something that should have passed.
Actually right
Report/grader issue
Pick every reason you think applies. Multiple boxes are allowed because one utterance can be wrong in more than one way. Use the note below for "Other" or extra detail.
Timing 0.5s
1 · TTS said no speech captured
2 · Card shown How much should I log for Grapes? I did not log them yet because the amounts were not clear.
3 · App data rows written No food/workout/list/note data row captured for this utterance.This does not mean "nothing anywhere was stored"; spoken/chat evidence is shown above, and some chat/session records may live outside this write-capture field. It means the report did not capture an app data-row write such as food_log_entries or workout_sets.
4 · UI did no screen events captured
5 · Seed/context data
Seeded starting state: none listed in the corpus row.
  • No DB snapshot captured for this utterance; the report cannot prove whether yesterday/personal data existed.
snapshot captured 2026-08-01T13:26:08.757Z
6 · Why this food No food-decision explanation was captured for this path.
7 · Lookup method No lookup method was captured for this path.
8 · Candidates → decider no candidate list captured (this path did not run DB resolution — e.g. provisional branded log)