Best fit
An institution deciding about many people at once, where appeals exist but do not change the rule. Public benefits is the pattern behind two of the five scored cases.
Public benefits and debt recovery →The standards bind a decision, not a model, so they apply in any sector. They do the most in six. This page names them, says what to bind and run first in each, and names five places where the framework does less.
On this page
An institution deciding about many people at once, where appeals exist but do not change the rule. Public benefits is the pattern behind two of the five scored cases.
Public benefits and debt recovery →A team shipping an agent or an automated decision whose code it controls. The record format is specified, and an exported record stream can be graded today.
AI agents that take actions →An agency or a health plan buying one of these systems from a vendor. It can cite clause numbers in the contract and require an evidence pack at each release.
Citing the standards in contracts →The framework earns its cost where four things are true.
When all four hold, the casebook pattern can recur: appeals are won one at a time while the rule that produced them keeps running. When one is missing, the framework does less. The boundaries below say what.
Each names what the system decides, the scored cases that came from it, the standards to bind, and the check to run first.
Whether a person gets a benefit, how much, and whether they owe it back.
Two of the five scored cases are here. In both, people contested their decisions for years while the rule that produced the errors kept running. The burden of disproof often falls on the claimant, and a wrong debt stands while they contest it.
The usual defense: "Each wrong decision was corrected on appeal." A corrected case is not a corrected rule. A claimant who cannot opt out is owed more correction, not less.
Whether a person gets credit, at what limit, and whether they can use the money or the account they already have.
These decisions run in real time, and many take effect before anyone could review them. The framework does not ask them to wait. It asks that each hold issue a receipt, run review clocks, and restore access when no fraud is confirmed.
The usual defense: "The fraud model is accurate." Accuracy is evidence for a hold. It is not permission to keep one open without a receipt or a clock.
Whether to send, refund, write, deploy, or approve, under authority someone delegated to it.
The tooling is furthest along here. Agent decisions are cheap and numerous, and a logging layer may not record them as decisions at all. STD-07 specifies the record each action leaves. STD-08 treats the agent's authority as a lease that expires. STD-09 treats a chain of agents as one delegation.
The usual defense: "The agent passed its evals." Passing evals shows what the agent can do. It still needs a grant that names its scope and expires.
Whether a treatment is authorized or a claim is paid, and how a patient is triaged.
Coverage decisions are frequent and each one is consequential, and a delay is itself a harm. Patients and clinicians already appeal them, so the casebook's question applies directly: does a won appeal change the rule?
The usual defense: "Patients can appeal." Appeals that win one case at a time, while the rule stays, show the rule is wrong. They do not show the system works.
Whether a person's post stays up, their account stays open, or their listing stays visible.
Enforcement runs at a volume no team could review case by case, and a wrong suspension can cut off a person's income or audience. Removal often has to be fast. The framework asks that it issue a receipt, bound its exceptions, and restore the content if the takedown is overturned.
The usual defense: "Users can appeal a takedown." An appeal is a remedy only if an overturned takedown restores the content, on a clock.
What grade a student receives, or whether a worker is accused, disciplined, or let go on the system's figures.
Two scored cases are here, and neither system was a machine-learning model. Ofqual's 2020 model was a statistical standardization. Horizon was branch accounting software. The standards apply anyway, because they bind the decision, not the component that made it.
The usual defense: "The system is robust." The Post Office told each defendant that of Horizon, while subpostmasters were held liable for the shortfalls it reported.
"Law in force" names rules that already ask for the same evidence. It is not legal advice, and meeting a standard here does not show that a system complies with any of them.
Five kinds of system where one of the four conditions fails, and what still applies.
No single recommendation changes a person's status, access, money, or risk, so there is no decision to contest. The framework applies where personalization sets a price, a credit offer, or who is shown a job.
The standards evaluate a delegation, not a model. A model can pass every benchmark and still be deployed under a grant nobody renews or can revoke. A model release has no one the standards can bind until someone deploys it.
When a named person decides each case, has the time and the information to disagree, and their disagreements are recorded, most of the framework's work is done. What remains is checking that those three stay true as volume grows.
A board that approves ten grants a year reads each one. Receipts, clocks, and conformance levels add cost there without catching what ordinary review would miss.
A build pipeline or an internal dashboard changes no one's status. The exception is an agent with write access to customer accounts or production records. That is a delegation, and the agents context applies.
A tool that checks evidence and a tool that scores answers are different instruments. So is a standard nobody has adopted yet.
The Record Conformance Checker reads an exported stream of STD-07 records, checks it against the schema, recomputes the hashes, and reports the conformance level the stream earns. The receipt and record formats are published as JSON Schemas that any validator can run in CI.
The Delegation Audit, the corrective capacity self-assessment, the Workload Modeler, the three validators, and the 60-second self-test score what a team enters about its system. None of them verifies it.
Every standard here is a draft, apart from STD-04, which is deprecated. None has force until an organization adopts it or a contract cites it. The case scores are this project's reading of what courts, inquiries, and regulators found.
Start with one workflow. The Delegation Audit takes it through six questions and names the actions nobody can ground in a grant.