Eval runner
Governance eval runner
Step through test cases one at a time. Score each one, capture evidence, and generate a report with aggregate score and grade.
In Phase 1, all scoring is manual. Each test case includes a prompt, system context, pass criteria, fail indicators, and a scoring rubric. The runner does not call any AI — it structures the review process and computes results from your manual scores.