Persona-Memory Evaluation Dashboard

Shows synthetic-eval quality and live-smoke health from the latest harness artifacts. Deterministic quality should stay green before merge; live smoke should pass or be explicitly skipped.

Synthetic Eval Action

Calls POST /api/run/eval, runs bun tests/eval.ts --report json, uses deterministic fixtures only (no network), and updates tests/.artifacts/eval-report.json.

Live Smoke Action

Calls POST /api/run/live, runs bun tests/run-live.ts --report json, attempts Anthropic smoke checks, and writes tests/.artifacts/live-report.json.

Step 1
Run
Execute deterministic gate first to refresh artifacts and assert baseline quality.
Step 2
Review
Inspect overall score, thresholds, and event/query outcomes to spot weak rows.
Jump To Review
Step 3
Curate
Label rows with rationale, assign score (1-5), and edit synthetic seeds when needed.
Jump To Curation
Step 4
Rebuild
Compile effective evalset so deterministic runs include your latest curation changes.

Run Activity

No run has been triggered in this browser session.
Click Run Synthetic Eval or Run Live Smoke to execute tests. This panel will show endpoint, script command, artifact writes, and outcome.

Overall Health

No data loaded.
Score 0 / 100
Load artifacts to compute a quality score.

Deterministic Eval Metrics

Waiting for eval-report.json

Threshold Checks

Direct pass/fail from synthetic eval checks

Live Smoke

Waiting for live-report.json

Counts and Throughput

Synthetic Dataset Inputs

Waiting for fixture rows.

Evalset Curation

Use this panel to hand-label rows and rebuild compiled evalset.
Annotation Editor
Select Mark Good/Mark Bad on a row to populate the editor.
Base Row Editor
Select Edit Row on a base row to update synthetic seed data.

Question to Answer Outcomes

Waiting for query outcomes from eval-report.json.

Event Outcome Rows

Waiting for event outcomes from eval-report.json.

Default artifact paths are loaded relative to this file: ../../tests/.artifacts/eval-report.json and ../../tests/.artifacts/live-report.json. If opening via file://, use file upload buttons or run bun scripts/serve-eval-dashboard.ts from the skill root for one-click eval runs.