The appeals kept winning. The system kept running.
That was Robodebt, Australia's automated welfare-debt scheme. We publish
open standards, public case scores, and diagnostics for one question:
when an automated decision system is hurting people, does evidence of
the harm reach someone who has to change the system?
The discipline that keeps automated decision systems answerable: when one is wrong, evidence of the harm reaches someone who has to change it.
Scope & applicability
What counts as a system under test?
If software takes an action, grants a benefit, denies an application,
or allocates work—and a human being has to live with the consequences
of an error—this is an initial scope filter. The full test also
requires adoption without independent review, a viable objection
route, and a named accountable institution.
The object being engineered is the delegation and its authority lease,
not model capabilities.
We reject "human-in-the-loop" when it means human liability sponging.
When upheld challenges cross a precommitted threshold, the automated
rule goes to its named owner for review, with revision or a reasoned
refusal recorded.
Closed feedback loops, not vague principles like "fairness" or
"transparency".
Spend 5 minutes running a self-test diagnostic, 1 hour evaluating an
agent against a Finite dual-ledger surge drill, or 1 sprint adopting
machine-readable delegation records.
Open standards, open data, and open benchmark fixtures licensed under
CC BY-SA 4.0.
The automated scheme launches. Debt notices go from about 20,000 a year to 20,000 a week. Every dated warning the record issued follows, with the
response each got.
“Income averaging cannot prove a debt.”
2014 · Legal advice
no response
“Notices do not explain how a debt was calculated.”
Apr 2017 · The Commonwealth Ombudsman
no response
“The debt is unlawful.”
— Dozens of times. Each ruling fixes one case. The scheme keeps running.
2017–19 · The Administrative Appeals Tribunal
no response
“A debt raised by averaging was not lawfully made.”
Nov 2019 · The Commonwealth
halt
“The scheme was unlawful from the outset.”
Jul 2023 · The Royal Commission
came after the halt
The rail is the count: rust dots are warnings the institution never
answered by changing the scheme.
Read the scored case →
The first three lines ran at Robodebt. The fourth is the line the
standards require.
Why the rulings did not stop it
Each ruling fixed one debt. None of them reached the rule.
The department treated each ruling as one person's outcome, not as
evidence against the scheme, so the rule that raised the debts never
had to answer them. The fix is a count: upheld challenges recorded
against the rule that produced them, and a set number that sends the
rule back to its owner.
A court, an inquiry, or a regulator established the facts of each. None
was stopped by the organization running it. Apple Card's issuer is the one
operator that changed its own process, after a regulator's investigation.
Each row: the six safeguards as a verdict barcode, and the time to halt as
the bar at the right edge — longest harm first.
The standards do not say who should hold power. They hold an institution to what it says about itself. Each claim below is set as a clause, with the exhibit an institution would have to produce.
§1
If the system is called democratic or legitimate
Exhibit A
— Show who is exposed to its decisions, who can challenge them, whose challenge must be answered by a date, and who can force a reconsideration or a halt.
Exhibit B
— Show that the figure counts the work the system pushes onto people with less power: the nurse fixing a scheduler's mistakes, the claimant proving a denial wrong.
An election can authorize a program. It does not give the person the program
decides about a way to contest that decision; standing does. For how these
standards relate to the GDPR, the EU AI Act, the NIST AI RMF, ISO/IEC 42001,
and the OECD AI Principles, see the
standards comparison.
60-second self-test
When yours is wrong, does anyone have to act?
Pick one automated decision system in your own organization: one that
decides benefits, loans, schedules, or fraud alerts. Answer one question
about each of its six safeguards. "Not sure" counts as no. The score
reflects your answers; it does not check the system.