Fit

Which systems this is for

The standards bind a decision, not a model, so they apply in any sector. They do the most in six. This page names them, says what to bind and run first in each, and names five places where the framework does less.

Best fit

An institution deciding about many people at once, where appeals exist but do not change the rule. Public benefits is the pattern behind two of the five scored cases.

Public benefits and debt recovery →

Furthest along

A team shipping an agent or an automated decision whose code it controls. The record format is specified, and an exported record stream can be graded today.

AI agents that take actions →

For a buyer

An agency or a health plan buying one of these systems from a vendor. It can cite clause numbers in the contract and require an evidence pack at each release.

Citing the standards in contracts →

The test

The framework earns its cost where four things are true.

  1. The decision changes one person's status, access, money, or risk.
  2. It runs at a volume nobody reviews case by case before it takes effect.
  3. The person can object, or should be able to, and the objections arrive somewhere the operator could read them.
  4. A named institution operates it and can be held to a clause by a contract, a regulator, or a court.

When all four hold, the casebook pattern can recur: appeals are won one at a time while the rule that produced them keeps running. When one is missing, the framework does less. The boundaries below say what.

Six contexts

Each names what the system decides, the scored cases that came from it, the standards to bind, and the check to run first.

Public benefits and debt recovery

Whether a person gets a benefit, how much, and whether they owe it back.

Two of the five scored cases are here. In both, people contested their decisions for years while the rule that produced the errors kept running. The burden of disproof often falls on the claimant, and a wrong debt stands while they contest it.

The usual defense: "Each wrong decision was corrected on appeal." A corrected case is not a corrected rule. A claimant who cannot opt out is owed more correction, not less.

Systems
Eligibility checks, overpayment and debt raising, fraud risk scores on claims, data matching across agencies.
Scored cases
Run first
VAL-01 Burden Modeler. Score how long the claim or appeal journey takes the claimant, and whether it is hard enough to amount to a denial.
Composite scenarios
Law in force
EU AI Act Annex III, point 5(a), makes high-risk the systems public authorities use to grant, reduce, revoke, or reclaim public assistance benefits, with obligations from 2 December 2027. GDPR Article 22 gives a right to contest a decision based solely on automated processing.
The person decided about can ask
In the EU, if a decision about you was made solely by automated processing, GDPR Article 22 lets you ask for a person to review it, give your view, and contest it. Article 15 lets you ask for meaningful information about the logic involved.

Credit, payments, and account holds

Whether a person gets credit, at what limit, and whether they can use the money or the account they already have.

These decisions run in real time, and many take effect before anyone could review them. The framework does not ask them to wait. It asks that each hold issue a receipt, run review clocks, and restore access when no fraud is confirmed.

The usual defense: "The fraud model is accurate." Accuracy is evidence for a hold. It is not permission to keep one open without a receipt or a clock.

Systems
Credit scoring and limit setting, fraud holds, account locks, payment blocks, refund and chargeback decisions.
Scored cases
Run first
Delegation Audit. Take one hold or limit workflow through six questions. See which actions nobody can ground in a grant, and whether each can be undone.
Composite scenarios
Law in force
In the US, ECOA and Regulation B require the specific principal reasons for an adverse credit action, and CFPB Circular 2022-03 says a complex model does not excuse a creditor from giving them. EU AI Act Annex III, point 5(b), makes creditworthiness assessment high-risk and excludes fraud detection.
The person decided about can ask
In the US, if you are refused credit or offered worse terms than you asked for, ECOA and Regulation B entitle you to the specific principal reasons. If the notice does not give them, you can ask within 60 days.

AI agents that take actions

Whether to send, refund, write, deploy, or approve, under authority someone delegated to it.

The tooling is furthest along here. Agent decisions are cheap and numerous, and a logging layer may not record them as decisions at all. STD-07 specifies the record each action leaves. STD-08 treats the agent's authority as a lease that expires. STD-09 treats a chain of agents as one delegation.

The usual defense: "The agent passed its evals." Passing evals shows what the agent can do. It still needs a grant that names its scope and expires.

Systems
Support and refund agents, agents with write access to accounts or records, chains of agents handing work to each other, typed decision models used as gates.
Scored cases
No public failure from this context has been scored yet. The refund example applies each clause to one agent.
Bind
Run first
Record Conformance Checker. Grade an exported stream of the agent's STD-07 records. If it emits none yet, start with the Delegation Audit.
Law in force
No statute treats agent actions as a category. An agent takes on the obligations of the decision it makes: a refund agent answers to consumer law, and an agent that decides benefit eligibility is high-risk under Annex III.

Health coverage and care decisions

Whether a treatment is authorized or a claim is paid, and how a patient is triaged.

Coverage decisions are frequent and each one is consequential, and a delay is itself a harm. Patients and clinicians already appeal them, so the casebook's question applies directly: does a won appeal change the rule?

The usual defense: "Patients can appeal." Appeals that win one case at a time, while the rule stays, show the rule is wrong. They do not show the system works.

Systems
Prior authorization and utilization review, claim denials, clinical risk scores that set a care pathway, triage.
Scored cases
No case from health care has been scored yet. One composite scenario, an appeal accepted without a remedy, is drawn from this setting.
Run first
VAL-03 Latency Audit. Check observed decision times against the declared deadline, such as 72 hours for an urgent request, and whether a person can be reached to escalate.
Composite scenarios
Law in force
From January 2026, CMS-0057-F requires Medicare Advantage plans and Medicaid and CHIP programs to decide urgent prior authorization requests within 72 hours and standard ones within seven calendar days, and to give a specific reason for each denial. EU AI Act Annex III, point 5(c), covers risk assessment and pricing in life and health insurance.
Commissioned help
Ethotechnics Studio does commissioned safety evaluation for healthcare AI: a safeguards review, a readiness sprint on one workflow, investor diligence on a healthcare AI deal, or a clinical AI safety evaluation against FDA and EU AI Act expectations. Ethotechnics Studio →
The person decided about can ask
In the US, if a Medicare Advantage plan or a Medicaid or CHIP program denies a prior authorization, it must give you a specific reason. It must decide within 72 hours for an urgent request and seven calendar days for a standard one.

Platform enforcement

Whether a person's post stays up, their account stays open, or their listing stays visible.

Enforcement runs at a volume no team could review case by case, and a wrong suspension can cut off a person's income or audience. Removal often has to be fast. The framework asks that it issue a receipt, bound its exceptions, and restore the content if the takedown is overturned.

The usual defense: "Users can appeal a takedown." An appeal is a remedy only if an overturned takedown restores the content, on a clock.

Systems
Content takedowns, account suspensions, seller and creator delisting, spam and abuse filters.
Scored cases
No platform case has been scored yet.
Run first
Corrective capacity self-assessment. Ask whether a successful appeal changes the filter or only restores one post.
Worked examples
Law in force
The EU Digital Services Act requires a statement of reasons for each restriction (Article 17), an internal complaint-handling system (Article 20), and access to out-of-court dispute settlement (Article 21).
The person decided about can ask
In the EU, a platform that removes your content or suspends your account must tell you why (DSA Article 17), take your complaint through its own system for at least six months after the decision (Article 20), and point you to an out-of-court dispute settlement body (Article 21).

Grading people and judging their work

What grade a student receives, or whether a worker is accused, disciplined, or let go on the system's figures.

Two scored cases are here, and neither system was a machine-learning model. Ofqual's 2020 model was a statistical standardization. Horizon was branch accounting software. The standards apply anyway, because they bind the decision, not the component that made it.

The usual defense: "The system is robust." The Post Office told each defendant that of Horizon, while subpostmasters were held liable for the shortfalls it reported.

Systems
Exam standardization and automated marking, productivity and shift scoring, accounting systems whose figures are used against the people who work in them.
Scored cases
Run first
Corrective capacity self-assessment. Ask whether handled exceptions ever changed the rule, or only the one grade or the one shortfall.
Worked examples
None yet.
Law in force
EU AI Act Annex III makes high-risk the systems that evaluate learning outcomes (point 3(b)) and those that inform promotion, termination, or task allocation, or that monitor and evaluate workers' performance (point 4(b)).

"Law in force" names rules that already ask for the same evidence. It is not legal advice, and meeting a standard here does not show that a system complies with any of them.

Where it does less

Five kinds of system where one of the four conditions fails, and what still applies.

Recommendations and personalization

No single recommendation changes a person's status, access, money, or risk, so there is no decision to contest. The framework applies where personalization sets a price, a credit offer, or who is shown a job.

Retail personalization →

Testing a model before anyone deploys it

The standards evaluate a delegation, not a model. A model can pass every benchmark and still be deployed under a grant nobody renews or can revoke. A model release has no one the standards can bind until someone deploys it.

What the evals can check →

Tools that inform a person who decides

When a named person decides each case, has the time and the information to disagree, and their disagreements are recorded, most of the framework's work is done. What remains is checking that those three stay true as volume grows.

Meaningful control →

Decisions a committee makes a few times a year

A board that approves ten grants a year reads each one. Receipts, clocks, and conformance levels add cost there without catching what ordinary review would miss.

Automation that touches no one outside the team

A build pipeline or an internal dashboard changes no one's status. The exception is an agent with write access to customer accounts or production records. That is a delegation, and the agents context applies.

AI agents that take actions →

What runs, what is self-report, what is proposed

A tool that checks evidence and a tool that scores answers are different instruments. So is a standard nobody has adopted yet.

Checks evidence

The Record Conformance Checker reads an exported stream of STD-07 records, checks it against the schema, recomputes the hashes, and reports the conformance level the stream earns. The receipt and record formats are published as JSON Schemas that any validator can run in CI.

Record Conformance Checker · Receipt schema

Organizes a self-report

The Delegation Audit, the corrective capacity self-assessment, the Workload Modeler, the three validators, and the 60-second self-test score what a team enters about its system. None of them verifies it.

Delegation Audit · All diagnostics

Proposes

Every standard here is a draft, apart from STD-04, which is deprecated. None has force until an organization adopts it or a contract cites it. The case scores are this project's reading of what courts, inquiries, and regulators found.

Standards and their status · How the cases are scored

If your system fits

Start with one workflow. The Delegation Audit takes it through six questions and names the actions nobody can ground in a grant.