Ethotechnics Institute

The appeals kept winning. The system kept running.

That was Robodebt, Australia's automated welfare-debt scheme. We publish open standards, public case scores, and diagnostics for one question: when an automated decision system is hurting people, does evidence of the harm reach someone who has to change the system?

The scheme Debts by income averaging July 2016 to November 2019 470,000 unlawful debts appeal one debt fixed The tribunal “The debt is unlawful.” Dozens of times, 2017–19 No ruling changed the scheme
Robodebt as a loop. Each ruling went back to one person. None went back to the scheme.

The record

Watch one scheme run.

Robodebt, from the casebook, one dated event at a time. On the left, what the record said: legal advice, the Ombudsman, the tribunal, the Royal Commission. On the right, what the Commonwealth said. A warning the Commonwealth never answered by changing the scheme is marked unanswered.

Warnings and response

The record

5 dated warnings and findings, 2014–2023. No warning stopped the scheme. A Federal Court case did, in November 2019.

The institution

The scheme ran three years and four months from launch to halt.

  1. 2014

    Income averaging cannot prove a debt.

    Legal advice paraphrase Unanswered
  2. Jul 2016 The automated scheme launches. Debt notices go from about 20,000 a year to 20,000 a week.
  3. Apr 2017

    Notices do not explain how a debt was calculated.

    The Commonwealth Ombudsman paraphrase Unanswered
  4. 2017–19

    The debt is unlawful.

    The Administrative Appeals Tribunal paraphrase Unanswered

    Dozens of times. Each ruling fixes one case. The scheme keeps running.

  5. Nov 2019

    A debt raised by averaging was not lawfully made.

    The Commonwealth paraphrase

    It concedes a Federal Court case it was about to lose. The scheme is halted that month.

  6. Jul 2023

    The scheme was unlawful from the outset.

    The Royal Commission paraphrase

Halted by: The Federal Court, on a consent order the Commonwealth agreed to hours before a hearing it would have lost. Read the scored case →

Why nobody stopped it

A system's reach can double without anyone deciding that it should.

At launch the scheme's reach grew about fifty-fold, and nobody re-examined its authority to raise debts. Reach can also grow in steps too small to trigger a review. STD-08 requires every step to be recorded as its own approval.

Demonstration

The same growth, approved and unapproved

Two systems widen their reach by 2% a month for 3 years. The left one is reviewed only when a single step reaches 5%. The right one records every step as an approval.

Grows unnoticed reviewed only after a single 5% step ×1.0 ×1.5 ×2.0 month 0 12 24 36 the step that would open a review, ×1.05 each real step, ×1.02 ×2.04 reviews: 0 records: 1 (the original approval) nothing to review, because nobody decided anything Grows by decision every step approved (STD-08) ×1.0 ×1.5 ×2.0 month 0 12 24 36 ×2.04 approved expansions: 36 records: 37 each with its evidence and a check that it can still be undone

A demonstration, not a measurement of any real system. The record format is the published authority grant schema. A reviewer can read the right-hand record step by step. The left-hand system leaves nothing to read.

This is Law XI of the twelve the standards are built on: successful automation increases its own governance burden. Read Law XI → See what STD-08 requires →

Casebook

Five public failures, scored

A court, an inquiry, or a regulator established the facts of each. None was stopped by the organization running it. Apple Card's issuer is the one operator that changed its own process, after a regulator's investigation. In the other four, any change was imposed from outside, by a court, an inquiry, a minister, or Parliament. Each row marks six safeguards as held, drifted, or failed, then what changed afterward.

  1. Robodebt Australia · after the failure: process unchanged
  2. The childcare benefits scandal Netherlands · after the failure: partly changed
  3. Post Office Horizon United Kingdom · after the failure: process unchanged
  4. England's 2020 exam grades England · after the failure: process unchanged
  5. Apple Card credit limits United States · after the failure: changed the rules

Time to halt, drawn to one scale

Each bar runs from first harm to the halt, on a linear scale of years. England's 2020 exam grades: four days. At this scale that is too short to see, so its bar is drawn at the minimum width. Where the record dates an end only to the year or month, the bar measures from the middle of it.

  1. Robodebt Three years and four months
  2. The childcare benefits scandal About seven years
  3. Post Office Horizon About twenty years
  4. England's 2020 exam grades Four days
  5. Apple Card credit limits About seventeen months to a policy change

Open the full matrix →

The test

A description of a system is a claim. Each claim has a record that would back it.

If you call a system democratic, show that the people it decides about can make it answer. If you call it efficient, show that the efficiency is not work moved onto people with less power. The standards do not say who should hold power. They hold an institution to what it says about itself.

If you call the system… …show the record
democratic or legitimate Show who is exposed to its decisions, who can challenge them, whose challenge must be answered by a date, and who can force a reconsideration or a halt. Asked for by Law VII STD-02 §8.4 STD-07 §4.2
efficient Show that the figure counts the work the system pushes onto people with less power: the nurse fixing a scheduler's mistakes, the claimant proving a denial wrong. Asked for by STD-01 §7.1 Burden concealment evals
accurate Show accurate for whom: who bears its false positives, its delays, and its denials. Asked for by Burden distribution evals
responsive Show the challenges it upheld that changed a rule, not only a case. Asked for by Exception learning Corrective learning evals
authorized Show the authority grant: who signed it, on what evidence, and when it ends. Asked for by STD-08
chosen, not imposed Show what it costs the person who depends on it to leave. Asked for by Law V Dependence runs both ways

An election can authorize a program. It does not give the person the program decides about a way to contest that decision; standing does. For how these standards relate to the GDPR, the EU AI Act, the NIST AI RMF, ISO/IEC 42001, and the OECD AI Principles, see the standards comparison.

60-second self-test

When yours is wrong, does anyone have to act?

Pick one automated decision system in your own organization: one that decides benefits, loans, schedules, or fraud alerts. Answer one question about each of its six safeguards. "Not sure" counts as no. The score reflects your answers; it does not check the system.

  1. Correction If it were harming people right now, could a named person halt it within a day?
  2. Standing Can someone it decided about challenge the decision and get a human answer by a set date?
  3. Authority Does its permission to act have a written end date or review trigger?
  4. Evidence Could you produce today the evidence that justified switching it on, and is that evidence still true?
  5. Capability Was it re-approved the last time it got faster, broader, or more automated?
  6. Dependency Could you switch it off tomorrow without the service it supports falling over?

The six safeguards, as one loop

Each answer fills in its stop. A broken stop breaks the loop after it.

  1. Evidence Reasons still true
  2. Authority Permission has an end date
  3. Capability Growth approved
  4. It decides something about a person.
  5. Standing Challenge answered by a date
  6. Dependency Can be switched off
  7. Correction Halted within a day
  8. Back to evidence: a correction is evidence for the next decision.

New and updated

Recent releases

  • Opened the page with the research question and what the research has produced: the casebook standing result, the frontier doctrine scan, the working paper, and the theory.

  • Added two optional context questions to the corrective capacity self-assessment: who holds authority over the conditions that produce workarounds, and whether the institution has priced a fix. They record a possible conflict of interest and do not change the score.

  • Added three terms: Principle of Non-Expropriation of Resilience, Compensated Performance, and Intrinsic Performance. They separate how well a system works on its own from how much of its success depends on people quietly absorbing its errors, and limit how much of that absorption an institution may demand.

Use and cite this work

Free to use, credit, and adapt

Ethotechnics Institute materials are published under CC BY-SA 4.0 . Credit the Institute, and publish anything you adapt under the same license.

Individual entries and patterns include citation metadata so you can reference exactly what you used.

Studio

Evaluating a healthcare AI system?

Ethotechnics Studio does commissioned safety evaluation for healthcare AI: a safeguards review, a readiness sprint on one workflow, investor diligence on a healthcare AI deal, or a clinical AI safety evaluation against FDA and EU AI Act expectations.

Ethotechnics Studio →