Test runs, clustered by corpus β each corpus's most recent run on top, its earlier runs (β³) nested underneath so you can watch that corpus improve over time. Clusters ordered by most-recent activity; any live run pins its corpus to the top. PASS = the app did what each utterance was supposed to do (log, clarify, deleteβ¦ per its intent); FAIL = it didn't. Click a run to open its full per-utterance detail.
Most recent test activity: just now (live run in progress)
Live DB strip unavailable. The rich drill-down report below is still intact.
Testing: LIVE β just now
Run: lane-stager canary 913FA249 Recent traces: 396 in the last 15 minutes Β· Account: 98ebe460 Β· Trace: 940d2e68-3e20-4096-afe5-9af4cf146a4e
Pass/Fail/Unverified = share of ALL rows. Pass rate = pass Γ· (pass + fail) only β it leaves out Unverified rows (not yet human-reviewed), so it can read a little different from the Pass % in the column next to it. Avg difficulty shows "β" until A1's per-utterance difficulty score lands (requested 2026-07-05). Don't compare pass rate across different corpora β one may simply be harder than another; compare a corpus against its own earlier runs instead. This page auto-refreshes while a run is live.