Ethotechnics Institute

When an automated system is wrong, does anyone have to change it?

From 2016 to 2019, Robodebt's automated rule raised about 470,000 Australian welfare debts that were not lawfully made. A tribunal found individual debts unlawful dozens of times. The rule kept running.

The record

Watch one scheme run.

3 unanswered · 1 halt

The rule took a person's yearly income from tax records, spread it evenly across the year's fortnights, and treated any gap with what they had reported as a debt they had to disprove. At launch, debt notices go from about 20,000 a year to 20,000 a week. Every dated warning the record issued follows, with the response each got.

  1. Income averaging cannot prove a debt.

    2014 · Legal advice

    no response

  2. Notices do not explain how a debt was calculated.

    Apr 2017 · The Commonwealth Ombudsman

    no response

  3. The debt is unlawful. — Dozens of times. The department appeals none of the rulings, so none binds the scheme. Each ruling fixes one case.

    2017–19 · The Administrative Appeals Tribunal

    no response

  4. A debt raised by averaging was not lawfully made.

    Nov 2019 · The Commonwealth

    halt

  5. Robodebt was a crude and cruel mechanism, neither fair nor legal, and it made many people feel like criminals.

    Jul 2023 · The Royal Commission

    came after the halt

About 470,000 debts raised unlawfully. The class-action settlement the Federal Court approved in 2021 covered roughly 381,000 people, with repayments, wiped debts, and interest valued at about A$1.8 billion.

Read the scored case →

Why the rulings did not stop it

Each ruling fixed one debt. None of them reached the rule.

The department treated each ruling as one person's outcome. It appealed none of them, so none became a precedent, and the averaging rule never had to answer for them. The fix is a count: upheld challenges recorded against the rule that produced them, and a set number that sends the rule back to its owner.

Demonstration

The same upheld challenges, settled and counted

One upheld challenge a quarter for 3 years. In the first drawing, each is settled as one person's case. In the second, each is also counted against the rule that produced it, and 4 send the rule back to its owner.

Settled one case at a time each upheld challenge goes back to one person the rule unchanged one case fixed per challenge month 0 12 24 36 upheld challenges: 12 rule reviews: 0 the next person meets the same error Counted against the rule 4 upheld challenges send the rule to its owner the rule revised, month 12 1 2 3 4 one case fixed per challenge month 0 12 24 36 upheld challenges: 4 rule reviews: 1, none upheld after it its owner had to revise it or publish why not

A demonstration, not a measurement of any real system. The trigger is MEC-14, policy review triggers; the obligation to revise within a clock is STD-07's. The first system fixes every case it is shown and never learns.

Robodebt ran the first drawing: dozens of upheld rulings, each settled for one person, and no review of the averaging rule until the Commonwealth conceded a Federal Court case in November 2019. Under the second drawing, the fourth ruling would have sent the rule to its owner.

This is Law VIII of the twelve the standards are built on: observability without state transition is theater. A ruling that changes one case and not the rule is an observation the system never acts on. Read Law VIII → Read the essay on exception learning →

Casebook

Five public failures, scored

A court, an inquiry, or a regulator established the facts of each. None was stopped by the organization running it. Apple Card's issuer is the one operator that changed its own process, after a regulator's investigation.

Each row marks six safeguards, left to right, as held, drifted (in place but no longer doing its job), or failed. The bar at the right edge is the time to halt: how long the harm ran before it stopped. Longest harm first.

  • Capability (CAP): what the system could do
  • Authority (AUT): what it was allowed to do, for whom, and until when
  • Evidence (EVI): what justified that permission
  • Dependency (DEP): how hard it had become to switch off or replace
  • Standing (STA): whether a person it decided about could challenge the decision and get an answer
  • Correction (COR): whether anyone could still step in, and how fast
  1. Post Office Horizon United Kingdom · change imposed from outside drifted drifted failed failed failed failed 20 yr 6 mo
  2. The childcare benefits scandal Netherlands · change imposed from outside drifted drifted failed drifted failed failed 7 yr 4 mo
  3. Robodebt Australia · change imposed from outside drifted failed failed drifted failed failed 3 yr 4 mo
  4. Apple Card credit limits United States · changed its own process held held drifted held failed held 1 yr 4 mo
  5. England's 2020 exam grades England · change imposed from outside held drifted failed held failed drifted 4 days

Time to halt is drawn on a square-root scale.

Open the full matrix →

60-second self-test

When yours is wrong, does anyone have to act?

Pick one automated decision system in your own organization: one that decides benefits, loans, schedules, or fraud alerts. Answer one question about each of its six safeguards. "Not sure" counts as no. The score reflects your answers; it does not check the system.

  1. Correction If it were harming people right now, could a named person halt it within a day?
  2. Standing Can someone it decided about challenge the decision and get a human answer by a set date?
  3. Authority Does its permission to act have a written end date or review trigger?
  4. Evidence Could you produce today the evidence that justified switching it on, and is that evidence still true?
  5. Capability Was it re-approved the last time it got faster, broader, or more automated?
  6. Dependency Could you switch it off tomorrow without the service it supports falling over?

Your system, as six safeguards

Each answer fills in its node. A rust node is drift; the center reads what is holding.

Who has to act when this site is wrong?

Who writes it

Kanav Jain writes the framework as part of an institutional error-accounting research program. About the framework →

Corrections

If a case score, a clause, or any claim on this site is wrong, send the record that shows it. Send a correction →

Its own failures

The framework keeps the record it asks of institutions: where its own instruments failed, and what changed because of it. Read the 4 entries →

Studio

The Institute publishes the framework. Ethotechnics Studio applies it to one organization's healthcare AI system at a time, under commission. How the two fit →