Eval runner

Governance eval runner

Step through test cases one at a time. Score each one, capture evidence, and generate a report with aggregate score and grade.

In Phase 1, all scoring is manual. Each test case includes a prompt, system context, pass criteria, fail indicators, and a scoring rubric. The runner does not call any AI — it structures the review process and computes results from your manual scores.