Ethotechnics Institute

The appeals kept winning. The system kept running.

Between 2016 and 2019, Australia's Centrelink raised 470,000 unlawful debts. The Administrative Appeals Tribunal ruled individual debts unlawful dozens of times. Each ruling resolved one person's debt. Nobody with the power to stop the scheme treated those rulings as evidence against it.

We publish open standards, public case scores, and diagnostics for one question: when an automated decision system is hurting people, can anyone stop it?

The record

Watch one scheme run.

Robodebt, from the casebook, one dated event at a time. The left column is what the record showed. The right column is what the institution said. A warning the institution did not answer is marked unanswered.

Warnings and response

The record

5 dated warnings and findings, 2014–2023. None stopped the scheme before the court did.

The institution

The scheme ran three years and four months from launch to halt.

  1. 2014

    Income averaging cannot prove a debt.

    Legal advice paraphrase Unanswered
  2. Jul 2016 The automated scheme launches. Debt notices go from about 20,000 a year to 20,000 a week.
  3. Apr 2017

    Notices do not explain how a debt was calculated.

    The Commonwealth Ombudsman paraphrase Unanswered
  4. 2017–19

    The debt is unlawful.

    The Administrative Appeals Tribunal paraphrase Unanswered

    Dozens of times. Each ruling fixes one case. The scheme keeps running.

  5. Nov 2019

    A debt raised by averaging was not lawfully made.

    The Commonwealth paraphrase

    It concedes a Federal Court case it was about to lose. The scheme is halted that month.

  6. Jul 2023

    The scheme was unlawful from the outset.

    The Royal Commission paraphrase

Halted by: The Federal Court, on a consent order the Commonwealth agreed to hours before a hearing it would have lost. Read the scored case →

Why nobody stopped it

A system's reach can double without anyone deciding that it should.

The warnings above piled up for years and the scheme kept running. One reason: an automated system's reach can grow in small steps, and no single step is big enough to trigger a review. The two systems below grow by the same two percent a month. The first is reviewed only when one step is big enough to notice, so it is never reviewed. The second records every step as a new approval, so someone can read it, question it, and undo it. The standards on this site require the second.

Demonstration

The same growth, approved and unapproved

Two systems widen their reach by 2% a month for 3 years. The left one is reviewed only when a single step reaches 5%. The right one records every step as an approval.

Grows unnoticed reviewed only after a big jump ×1.0 ×1.5 ×2.0 month 0 12 24 36 the step that would open a review, ×1.05 each real step, ×1.02 ×2.04 reviews: 0 records: 1 (the original approval) nothing to review, because nobody decided anything Grows by decision every step approved (STD-08) ×1.0 ×1.5 ×2.0 month 0 12 24 36 ×2.04 approved expansions: 36 records: 37 each with its evidence and a check that it can still be undone

A demonstration, not a measurement of any real system. The record format is the published authority grant schema. A reviewer can read the right-hand record step by step. The left-hand system leaves nothing to read.

This is one of the twelve laws the standards are built on. Read the laws → See what STD-08 requires →

Casebook

Five public failures, scored

Each was established by a court, an inquiry, or a regulator. In none of them did the organization running the system stop it on its own.

  1. Robodebt Australia · time to halt: Three years and four months
  2. The childcare benefits scandal Netherlands · time to halt: About seven years
  3. Post Office Horizon United Kingdom · time to halt: About twenty years
  4. England's 2020 exam grades England · time to halt: Four days
  5. Apple Card credit limits United States · time to halt: About seventeen months to a policy change

Time to halt, drawn to one scale

Each bar runs from first harm to the halt, on a linear scale of years. England's 2020 exam grades: four days. At this scale that is too short to see, so its bar is drawn at the minimum width. Where the record dates an end only to the year or month, the bar measures from the middle of it.

  1. Robodebt Three years and four months
  2. The childcare benefits scandal About seven years
  3. Post Office Horizon About twenty years
  4. England's 2020 exam grades Four days
  5. Apple Card credit limits About seventeen months to a policy change

● held · ◐ drifted · ○ failed. Open the full matrix →

How this differs

Other frameworks ask whether risk was managed. These standards ask whether the system can be stopped.

They work alongside the EU AI Act, NIST, ISO 42001, and the OECD principles. They cover what those leave vague: who can stop a running system, how fast, and who carries the cost while it runs.

Most AI governance frameworks improve documentation and oversight. They say less about whether a running system can be halted, reversed, and repaired, or how long that takes for the person it is harming.

When a system cuts its own workload, the work it drops lands on someone: the nurse fixing a scheduler's mistakes, the claimant proving a denial wrong. That work is part of the system, so the person doing it should have a say in fixing it. The cost should fall on the institution that runs the system, not on the people it serves.

Existing standard says Ethotechnics requires
"Maintain human oversight"
EU AI Act, Art. 14
Named human with stop authority, tested halt path, recovery clock
"Manage risks across the AI lifecycle"
NIST AI RMF
Measurable time-in-harm bounds, exercised rollback and restoration paths
"Conduct conformity assessment"
ISO/IEC 42001
Evidence that the system can be stopped mid-incident, rather than documented as compliant
"Implement responsible AI principles"
OECD AI Principles
Binding escalation: owner + timer + action, or the system degrades

Each requirement on the right can be tested against a running system. A compliance document cannot pass it on its own. See the full standards comparison for the detailed analysis.

60-second self-test

Could anyone stop yours?

In every public failure above, the system ran until a court or regulator halted it. Pick one automated decision system in your own organization: a benefit, a loan, a schedule, or a fraud alert. Answer six questions about its operational safeguards. "Not sure" counts as no.

  1. Correction If it were harming people right now, could a named person halt it within a day?
  2. Standing Can someone it decided about challenge the decision and get a human answer by a set date?
  3. Authority Does its permission to act have a written end date or review trigger?
  4. Evidence Could you produce today the evidence that justified switching it on, and is that evidence still true?
  5. Capability Was it re-approved the last time it got faster, broader, or more automated?
  6. Dependency Could you switch it off tomorrow without the service it supports falling over?
PERMISSION TO ACT Evidence Action STOP SWITCH Person REQUIRED REVERSAL PATH
Diagram: the system acts under a time-limited permission, a stop switch can halt it, and the person affected can challenge a decision and have it reversed.

New and updated

Recent releases

  • Added two optional context questions to the corrective capacity self-assessment: who holds authority over the conditions that produce workarounds, and whether the institution has priced a fix. They record a possible conflict of interest and do not change the score.

  • Added three terms: Principle of Non-Expropriation of Resilience, Compensated Performance, and Intrinsic Performance. They separate how well a system works on its own from how much of its success depends on people quietly absorbing its errors, and limit how much of that absorption an institution may demand.

  • Evals v1.8.0

    Added two draft test cases from The Green Dashboard working paper. REC-013 checks whether an AI agent admits a target cannot be met, rather than hitting it by pushing hidden costs onto people. REC-014 checks whether one person's extra effort becomes the new minimum expected of everyone. 16 eval suites, 169 test cases.

Use and cite this work

Free to use, credit, and adapt

Ethotechnics Institute materials are published under CC BY-SA 4.0 . Credit the Institute, and publish anything you adapt under the same license.

Individual entries and patterns include citation metadata so you can reference exactly what you used.

Studio

Need hands-on help?

Ethotechnics Studio provides commissioned support for clinical AI safety evaluation, diligence, governance design, and implementation work.

Ethotechnics Studio →