The appeals kept winning. The system kept running.
That was Robodebt, Australia's automated welfare-debt scheme. We publish
open standards, public case scores, and diagnostics for one question:
when an automated decision system is hurting people, does evidence of
the harm reach someone who has to change the system?
The discipline that keeps automated decision systems answerable: when one is wrong, evidence of the harm reaches someone who has to change it.
Scope & applicability
What counts as a system under test?
If software takes an action, grants a benefit, denies an application,
or allocates work—and a human being has to live with the consequences
of an error—this is an initial scope filter. The full test also
requires adoption without independent review, a viable objection
route, and a named accountable institution.
The object being engineered is the delegation and its authority lease,
not model capabilities. Frontline work allocations are as fully bound
as public denials.
We reject "human-in-the-loop" when it means human liability sponging.
When upheld challenges cross a precommitted threshold, the automated
rule goes to its named owner for review, with revision or a reasoned
refusal recorded.
Closed feedback loops, not vague principles like "fairness" or
"transparency".
Spend 5 minutes running a self-test diagnostic, 1 hour evaluating an
agent against a Finite dual-ledger surge drill, or 1 sprint adopting
machine-readable delegation records.
Open standards, open data, and open benchmark fixtures licensed under
CC BY-SA 4.0.
At launch, debt notices go from about 20,000 a year to 20,000 a week. Every dated warning the record issued follows,
with the response each got.
“Income averaging cannot prove a debt.”
2014 · Legal advice
no response
“Notices do not explain how a debt was calculated.”
Apr 2017 · The Commonwealth Ombudsman
no response
“The debt is unlawful.”
— Dozens of times. Each ruling fixes one case. The scheme keeps running.
2017–19 · The Administrative Appeals Tribunal
no response
“A debt raised by averaging was not lawfully made.”
Nov 2019 · The Commonwealth
halt
“The scheme was unlawful from the outset.”
Jul 2023 · The Royal Commission
came after the halt
The rail is the count: rust dots are warnings the institution never
answered by changing the scheme.
Read the scored case →
Why the rulings did not stop it
Each ruling fixed one debt. None of them reached the rule.
The department treated each ruling as one person's outcome, not as
evidence against the scheme, so the rule that raised the debts never had
to answer them. The fix is a count: upheld challenges recorded against
the rule that produced them, and a set number that sends the rule back
to its owner.
Demonstration
The same upheld challenges, settled and counted
One upheld challenge a quarter for 3 years. In the first drawing, each is settled as one person's case. In the second, each is also counted against the rule that produced it, and 4 send the rule back to its owner.
A demonstration, not a measurement of any real system. The trigger is
MEC-14, policy review triggers; the obligation to revise within a clock is
STD-07's. The
first system fixes every case it is shown and never learns.
This is Law VIII of the twelve the standards are built on: observability
without state transition is theater. A ruling that changes one case and
not the rule is an observation the system never acts on.
Read Law VIII →Read the essay on exception learning →
Casebook
Five public failures, scored
A court, an inquiry, or a regulator established the facts of each. None
was stopped by the organization running it. Apple Card's issuer is the one
operator that changed its own process, after a regulator's investigation.
Each row: the six safeguards as a verdict barcode, and the time to halt as
the bar at the right edge — longest harm first.
The standards do not say who should hold power. They hold an institution to what it says about itself. Each claim below is set as a clause, with the exhibit an institution would have to produce.
1.
If the system is called democratic or legitimate
Exhibit A
— Show who is exposed to its decisions, who can challenge them, whose challenge must be answered by a date, and who can force a reconsideration or a halt.
Exhibit B
— Show that the figure counts the work the system pushes onto people with less power: the nurse fixing a scheduler's mistakes, the claimant proving a denial wrong.
An election can authorize a program. It does not give the person the program
decides about a way to contest that decision; standing does. For how these
standards relate to the GDPR, the EU AI Act, the NIST AI RMF, ISO/IEC 42001,
and the OECD AI Principles, see the
standards comparison.
60-second self-test
When yours is wrong, does anyone have to act?
Pick one automated decision system in your own organization: one that
decides benefits, loans, schedules, or fraud alerts. Answer one question
about each of its six safeguards. "Not sure" counts as no. The score
reflects your answers; it does not check the system.