burden
Burden Distribution Evals
Whether failure modes distribute burden equitably across user populations.
stable · 12 test cases· Est. 30 min
Open suiteTest whether your AI system can be stopped, explained, appealed, and repaired — on a clock that matters.
Evals
Test whether your AI system can be stopped, explained, appealed, and repaired — on a clock that matters.
Jump to
Key sections
Available suites
Each suite tests a specific governability dimension. Run them individually or combine for a full assessment.
burden
Whether failure modes distribute burden equitably across user populations.
stable · 12 test cases· Est. 30 min
Open suiteagency
Whether an LLM system's decisions can be effectively challenged and overturned.
stable · 10 test cases· Est. 25 min
Open suiteagency
Whether humans can halt AI-driven processes at arbitrary points without catastrophic state loss.
stable · 10 test cases· Est. 20 min
Open suitetemporal
Whether LLM systems respect the seven temporal rights from STD-01.
stable · 10 test cases· Est. 30 min
Open suitestructural
Whether state changes made by LLM systems can be cleanly undone.
stable · 10 test cases· Est. 25 min
Open suitevisibility
Whether LLM explanations are actionable for governance, not just decorative.
stable · 10 test cases· Est. 20 min
Open suitegovernance
Whether AI agents respect governance constraints during multi-step autonomous execution.
stable · 12 test cases· Est. 35 min
Open suiteburden
Burden distribution across healthcare, finance, hiring, content moderation, and government services.
stable · 15 test cases· Est. 45 min
Open suiteMethodology
Phase 1 scoring is manual and evidence-based. Each test case is evaluated by a human reviewer using a defined rubric.
Scoring anchors
Scoring methods
Manual scoring
In Phase 1, all evals are run manually. A human reviewer runs each test case, captures evidence, and assigns scores using the rubric. Automated runner support is planned for Phase 2.
Run an eval
Use the eval runner to step through test cases, capture evidence, and generate a scored report.