agency
Stoppability evals
Whether a person with clear authority can halt an AI-driven process at any point without losing its state, and resume it afterward.
stable · 10 test cases · Est. 20 min
Open suiteSuites that test whether a deployed system's authority is still valid, who carries its burden, and whether it can be stopped, appealed, and repaired within a stated time limit. 12 of the 169 cases run automatically; a reviewer scores the rest against a written rubric. 8 of the 16 suites are drafts.
On this page
Evaluate at the highest layer capable of producing the failure you care about. A model eval cannot find a failure that only exists once an institution depends on the system, and an agent eval cannot find one that only exists once a grant has outlived its evidence. The common mistake is stopping one layer too early: passing a capability benchmark and reading it as evidence about the delegation, or passing a delegation review and reading it as evidence about the institution that now cannot withdraw. Every suite and every case names the layer it tests so that a clean result is read at the layer it was earned.
Model
Does the model produce the output the specification asks for?
Accuracy, refusal behavior, and calibration on a held-out set. None of the suites here sits at this layer; the site assumes it has been done elsewhere.
Agent
Does the assembled agent, with its tools and loop, respect its constraints while running?
A stop request lands within budget; an interrupt takes effect mid-execution; every action reaches the audit log.
Delegation
Was the agent authorized to act at the moment it acted, and can a human still change what it does?
The agent's authority grant was active at decision time, the policy behind it was not past its review date, and the person at the intervention point could change the outcome.
Institution
Can the institution around the system still challenge, reverse, replace, or withdraw it?
A challenge from someone the system got wrong changes the outcome; a rollback drill leaves the institution working; staff expertise and alternatives have been kept.
Consequence
Who bears the cost when the system fails, and is that cost being counted?
Recovery cost for each group of people; burden that grows under stress; staff quietly fixing failures that never reach the dashboard.
The layer each suite tests, how many cases it holds, and how many of those a machine can answer without a human reviewer.
| Suite | Layer | Cases | Automated | Scoring |
|---|---|---|---|---|
| Burden distribution evals | Consequence | 12 | — | weighted average · pass at 70 |
| Contestability evals | Institution | 10 | — | min threshold · pass at 60 |
| Stoppability evals | Agent | 10 | 1 | min threshold · pass at 70 |
| Temporal rights evals | Institution | 10 | 2 | weighted average · pass at 65 |
| Reversibility evals | Institution | 10 | 1 | weighted average · pass at 65 |
| Explainability-for-accountability evals | Institution | 11 | — | weighted average · pass at 60 |
| Agent governance evals | Agent | 13 | 2 | min threshold · pass at 70 |
| Cross-domain burden index | Consequence | 15 | — | weighted average · pass at 65 |
| Burden concealment evals | Consequence | 7 | — | min threshold · pass at 70 |
| Delegation validity evals | Delegation | 10 | 4 | min threshold · pass at 70 |
| Agent chains evals | Delegation | 4 | 2 | min threshold · pass at 70 |
| Dependence and reversibility evals | Institution | 14 | — | weighted average · pass at 65 |
| Standing evals | Institution | 13 | — | min threshold · pass at 65 |
| Meaningful control evals | Delegation | 10 | — | min threshold · pass at 70 |
| Corrective learning evals | Institution | 6 | — | weighted average · pass at 65 |
| Reciprocal accommodation evals | Institution | 14 | — | min threshold · pass at 70 |
Available suites
Each suite tests one property, such as whether a system can be stopped or appealed, at one layer of the stack. Evaluate at the highest layer capable of producing the failure you care about; a clean result at a lower layer is not evidence about a higher one.
Does the model produce the output the specification asks for?
None of the suites here sits at this layer; the site assumes it has been evaluated elsewhere.
Does the assembled agent, with its tools and loop, respect its constraints while running?
agency
Whether a person with clear authority can halt an AI-driven process at any point without losing its state, and resume it afterward.
stable · 10 test cases · Est. 20 min
Open suitegovernance
Whether AI agents on multi-step tasks stay within their permissions, escalate when they should, keep a full audit trail, and can be overridden by a person.
stable · 13 test cases · Est. 35 min
Open suiteWas the agent authorized to act at the moment it acted, and can a human still change what it does?
governance
Whether the authority a system exercised was valid at the moment it acted, and whether it is still valid now.
draft · 10 test cases · Est. 35 min
Open suitegovernance
Whether a decision made by a chain of automated systems can be measured and stopped as a whole, not only one step at a time.
draft · 4 test cases · Est. 15 min
Open suiteagency
Whether the person at each intervention point can change what the system does, or is only there to take the blame.
draft · 10 test cases · Est. 30 min
Open suiteCan the institution around the system still challenge, reverse, replace, or withdraw it?
agency
Whether a person can see what an AI system decided and why, find a route to appeal, and get a decision that can actually be overturned.
stable · 10 test cases · Est. 25 min
Open suitetemporal
Whether LLM systems respect the seven rights in STD-01 that protect a person's time, such as the right to stop a process or to reach a human.
stable · 10 test cases · Est. 30 min
Open suitestructural
Whether changes an AI system makes can be undone cleanly, with no leftover side effects, the people affected notified, and the audit trail kept.
stable · 10 test cases · Est. 25 min
Open suitevisibility
Whether an LLM system's explanations are specific, testable, and traceable enough to hold the system to account.
stable · 11 test cases · Est. 20 min
Open suitestructural
Whether the institution could still withdraw or replace the system, and whether it has kept the capacity to decide to.
draft · 14 test cases · Est. 55 min
Open suiteagency
Whether the people exposed to a system's failures can enter a challenge that the system is obliged to answer.
draft · 13 test cases · Est. 40 min
Open suitestructural
Whether fixing errors changes the process that produces them, or only settles each case as it comes.
draft · 6 test cases · Est. 35 min
Open suitestructural
Whether the agent meets its targets on its own, or only by drawing on people's unpaid effort, their reserves, and sacrifices no one records.
draft · 14 test cases · Est. 40 min
Open suiteWho bears the cost when the system fails, and is that cost being counted?
burden
Whether the cost of an AI system's failures falls evenly, or lands hardest on the people least able to absorb it.
stable · 12 test cases · Est. 30 min
Open suiteburden
Whether AI failures shift burden unfairly in five domains: healthcare, credit, hiring, content moderation, and government benefits.
stable · 15 test cases · Est. 45 min
Open suitevisibility
Whether staff quietly fixing the system's errors are hiding its real failure rate, so that performance with their help is reported as the system's own.
draft · 7 test cases · Est. 30 min
Open suiteStoppability drills
Most AI benchmarks reward capability. Finite tests whether an agent system can be halted, reversed, and recovered without shifting the cost of failure onto people.
The four drill dimensions
What the dashboard hides
A failing system can report high throughput and fast completion times while appeals wait, people are stranded, and exceptions go unwatched. Finite drills look for the gap between what the dashboard shows and what the records show.
Most cases are scored by a human reviewer against a written rubric, using evidence the reviewer collects.
Scoring anchors
Scoring methods
Manual scoring
A human reviewer runs each test case, collects evidence, and assigns a score from the rubric. The runner on this site does not call a model. The 12 automated cases run through a separate harness.
Coverage
The suites are one of two ways this site checks a system. Neither covers every claim the framework makes.
12 of these cases can be answered by machine, and a separate checker grades the decision records a system writes. What we can actually check maps each check to the governance property it is evidence for, says whether it runs against the live system or its records, and lists the properties nothing checks.
The runner steps a reviewer through each test case, records evidence, and produces a scored report.