Ethotechnics Institute
When an automated system is wrong, does anyone have to change it?
From 2016 to 2019, Robodebt's automated rule raised about 470,000 Australian welfare debts that were not lawfully made. A tribunal found individual debts unlawful dozens of times. The rule kept running.
The record
Watch one scheme run.
3 unanswered · 1 halt
The rule took a person's yearly income from tax records, spread it evenly across the year's fortnights, and treated any gap with what they had reported as a debt they had to disprove. At launch, debt notices go from about 20,000 a year to 20,000 a week. Every dated warning the record issued follows, with the response each got.
-
Income averaging cannot prove a debt.
2014 · Legal advice
no response
-
Notices do not explain how a debt was calculated.
Apr 2017 · The Commonwealth Ombudsman
no response
-
The debt is unlawful. — Dozens of times. The department appeals none of the rulings, so none binds the scheme. Each ruling fixes one case.
2017–19 · The Administrative Appeals Tribunal
no response
-
A debt raised by averaging was not lawfully made.
Nov 2019 · The Commonwealth
halt
-
Robodebt was a crude and cruel mechanism, neither fair nor legal, and it made many people feel like criminals.
Jul 2023 · The Royal Commission
came after the halt
About 470,000 debts raised unlawfully. The class-action settlement the Federal Court approved in 2021 covered roughly 381,000 people, with repayments, wiped debts, and interest valued at about A$1.8 billion.
Why the rulings did not stop it
Each ruling fixed one debt. None of them reached the rule.
The department treated each ruling as one person's outcome. It appealed none of them, so none became a precedent, and the averaging rule never had to answer for them. The fix is a count: upheld challenges recorded against the rule that produced them, and a set number that sends the rule back to its owner.
Demonstration
The same upheld challenges, settled and counted
One upheld challenge a quarter for 3 years. In the first drawing, each is settled as one person's case. In the second, each is also counted against the rule that produced it, and 4 send the rule back to its owner.
A demonstration, not a measurement of any real system. The trigger is MEC-14, policy review triggers; the obligation to revise within a clock is STD-07's. The first system fixes every case it is shown and never learns.
Robodebt ran the first drawing: dozens of upheld rulings, each settled for one person, and no review of the averaging rule until the Commonwealth conceded a Federal Court case in November 2019. Under the second drawing, the fourth ruling would have sent the rule to its owner.
This is Law VIII of the twelve the standards are built on: observability without state transition is theater. A ruling that changes one case and not the rule is an observation the system never acts on. Read Law VIII → Read the essay on exception learning →
Casebook
Five public failures, scored
A court, an inquiry, or a regulator established the facts of each. None was stopped by the organization running it. Apple Card's issuer is the one operator that changed its own process, after a regulator's investigation.
Each row marks six safeguards, left to right, as held, drifted (in place but no longer doing its job), or failed. The bar at the right edge is the time to halt: how long the harm ran before it stopped. Longest harm first.
- Capability (CAP): what the system could do
- Authority (AUT): what it was allowed to do, for whom, and until when
- Evidence (EVI): what justified that permission
- Dependency (DEP): how hard it had become to switch off or replace
- Standing (STA): whether a person it decided about could challenge the decision and get an answer
- Correction (COR): whether anyone could still step in, and how fast
- Post Office Horizon drifted drifted failed failed failed failed 20 yr 6 mo
- The childcare benefits scandal drifted drifted failed drifted failed failed 7 yr 4 mo
- Robodebt drifted failed failed drifted failed failed 3 yr 4 mo
- Apple Card credit limits held held drifted held failed held 1 yr 4 mo
- England's 2020 exam grades held drifted failed held failed drifted 4 days
Time to halt is drawn on a square-root scale.
60-second self-test
When yours is wrong, does anyone have to act?
Pick one automated decision system in your own organization: one that decides benefits, loans, schedules, or fraud alerts. Answer one question about each of its six safeguards. "Not sure" counts as no. The score reflects your answers; it does not check the system.
Your system, as six safeguards
Each answer fills in its node. A rust node is drift; the center reads what is holding.
Start from where you are
Five ways in.
- I build these systems The checks a system has to pass before it ships, and the tools that run them.
- I audit or regulate them Citable clauses, scored public cases, and the evidence each clause needs.
- I decide whether to deploy them What five public failures cost, and what to check before you commit to a deployment.
- A system decided something about me Three checks: can this decision be challenged, can the appeal change the outcome, and who holds power here?
- I study these systems What the framework adds to theory you already work with, what it asks of the mechanisms it competes with, and where the argument is open to dispute.
Not sure the framework applies to your system? See which systems it is for →
Who has to act when this site is wrong?
Who writes it
Kanav Jain writes the framework as part of an institutional error-accounting research program. About the framework →
Corrections
If a case score, a clause, or any claim on this site is wrong, send the record that shows it. Send a correction →
Its own failures
The framework keeps the record it asks of institutions: where its own instruments failed, and what changed because of it. Read the 4 entries →
Studio
The Institute publishes the framework. Ethotechnics Studio applies it to one organization's healthcare AI system at a time, under commission. How the two fit →