Ethotechnics Institute

The appeals kept winning. The system kept running.

That was Robodebt, Australia's automated welfare-debt scheme. We publish open standards, public case scores, and diagnostics for one question: when an automated decision system is hurting people, does evidence of the harm reach someone who has to change the system?

The record The institution
2014 The department is advised in writing that income averaging cannot prove a debt. no response
Apr 2017 The Ombudsman reports that notices do not explain how a debt was calculated. no response
2017–19 A tribunal rules individual debts unlawful dozens of times. no response
Every warning the record issued before the halt, and the response it got. The empty column is the finding.

Three things this site answers before you begin

The discipline that keeps automated decision systems answerable: when one is wrong, evidence of the harm reaches someone who has to change it.

Scope & applicability

What counts as a system under test?

If software takes an action, grants a benefit, denies an application, or allocates work—and a human being has to live with the consequences of an error—this is an initial scope filter. The full test also requires adoption without independent review, a viable objection route, and a named accountable institution.

The object being engineered is the delegation and its authority lease, not model capabilities.

Explore covered use cases →

The mechanism

How this differs from "AI Ethics"

We reject "human-in-the-loop" when it means human liability sponging. When upheld challenges cross a precommitted threshold, the automated rule goes to its named owner for review, with revision or a reasoned refusal recorded.

Closed feedback loops, not vague principles like "fairness" or "transparency".

See the 7-stage governed chain →

Adoption ladder

Start by time available

Spend 5 minutes running a self-test diagnostic, 1 hour evaluating an agent against a Finite dual-ledger surge drill, or 1 sprint adopting machine-readable delegation records.

Open standards, open data, and open benchmark fixtures licensed under CC BY-SA 4.0.

Pick your entry path →

The record

Watch one scheme run.

3 unanswered · 1 halt

The automated scheme launches. Debt notices go from about 20,000 a year to 20,000 a week. Every dated warning the record issued follows, with the response each got.

  1. “Income averaging cannot prove a debt.”

    2014 · Legal advice

    no response

  2. “Notices do not explain how a debt was calculated.”

    Apr 2017 · The Commonwealth Ombudsman

    no response

  3. “The debt is unlawful.” — Dozens of times. Each ruling fixes one case. The scheme keeps running.

    2017–19 · The Administrative Appeals Tribunal

    no response

  4. “A debt raised by averaging was not lawfully made.”

    Nov 2019 · The Commonwealth

    halt

  5. “The scheme was unlawful from the outset.”

    Jul 2023 · The Royal Commission

    came after the halt

The rail is the count: rust dots are warnings the institution never answered by changing the scheme. Read the scored case →

The scheme The person debts appeal ruling — back to one person The rule ✕ the ruling never reaches the rule
The first three lines ran at Robodebt. The fourth is the line the standards require.

Why the rulings did not stop it

Each ruling fixed one debt. None of them reached the rule.

The department treated each ruling as one person's outcome, not as evidence against the scheme, so the rule that raised the debts never had to answer them. The fix is a count: upheld challenges recorded against the rule that produced them, and a set number that sends the rule back to its owner.

This is Law VIII of the twelve the standards are built on: observability without state transition is theater. Read Law VIII → Read the essay on exception learning →

Casebook

Five public failures, scored

A court, an inquiry, or a regulator established the facts of each. None was stopped by the organization running it. Apple Card's issuer is the one operator that changed its own process, after a regulator's investigation. Each row: the six safeguards as a verdict barcode, and the time to halt as the bar at the right edge — longest harm first.

  1. Post Office Horizon United Kingdom · change imposed from outside drifted drifted failed failed failed failed 20y 6m
  2. The childcare benefits scandal Netherlands · change imposed from outside drifted drifted failed drifted failed failed 7y 4m
  3. Robodebt Australia · change imposed from outside drifted failed failed drifted failed failed 3y 4m
  4. Apple Card credit limits United States · changed its own process held held drifted held failed held 1y 4m
  5. England's 2020 exam grades England · change imposed from outside held drifted failed held failed drifted 4 days

Sorted by time to halt, longest first. England's 2020 exam grades: 4 days — at a square-root scale that is a sliver, but the label carries the figure.

Open the full matrix →

For scholars and governance thinkers

Every corrective mechanism itself changes under delegation.

Markets

After ten years of success, does competition still hold corrective power?

Law

After ten years of success, does the ruling still reach the rule?

Management

After ten years of success, does the manager still hold authority to correct?

Technology

After ten years, does the tool still let people correct the tool?

This framework

The same question, asked of its own instruments.

Shared values, new instruments → · The real disagreement → · Scholarly crossings →

The test

A description of a system is a claim. Each claim has a record that would back it.

The standards do not say who should hold power. They hold an institution to what it says about itself. Each claim below is set as a clause, with the exhibit an institution would have to produce.

  1. §1 If the system is called democratic or legitimate

    Exhibit A — Show who is exposed to its decisions, who can challenge them, whose challenge must be answered by a date, and who can force a reconsideration or a halt.

    Asked for by Law VII · STD-02 §8.4 · STD-07 §4.2

  2. §2 If the system is called efficient

    Exhibit B — Show that the figure counts the work the system pushes onto people with less power: the nurse fixing a scheduler's mistakes, the claimant proving a denial wrong.

    Asked for by STD-01 §7.1 · Burden concealment evals

  3. §3 If the system is called accurate

    Exhibit C — Show accurate for whom: who bears its false positives, its delays, and its denials.

    Asked for by Burden distribution evals

  4. §4 If the system is called responsive

    Exhibit D — Show the challenges it upheld that changed a rule, not only a case.

    Asked for by Exception learning · Corrective learning evals

  5. §5 If the system is called authorized

    Exhibit E — Show the authority grant: who signed it, on what evidence, and when it ends.

    Asked for by STD-08

  6. §6 If the system is called chosen, not imposed

    Exhibit F — Show what it costs the person who depends on it to leave.

    Asked for by Law V · Dependence runs both ways

An election can authorize a program. It does not give the person the program decides about a way to contest that decision; standing does. For how these standards relate to the GDPR, the EU AI Act, the NIST AI RMF, ISO/IEC 42001, and the OECD AI Principles, see the standards comparison.

60-second self-test

When yours is wrong, does anyone have to act?

Pick one automated decision system in your own organization: one that decides benefits, loans, schedules, or fraud alerts. Answer one question about each of its six safeguards. "Not sure" counts as no. The score reflects your answers; it does not check the system.

  1. Correction If it were harming people right now, could a named person halt it within a day?
  2. Standing Can someone it decided about challenge the decision and get a human answer by a set date?
  3. Authority Does its permission to act have a written end date or review trigger?
  4. Evidence Could you produce today the evidence that justified switching it on, and is that evidence still true?
  5. Capability Was it re-approved the last time it got faster, broader, or more automated?
  6. Dependency Could you switch it off tomorrow without the service it supports falling over?

Your system, as six safeguards

Each answer fills in its node. A rust node is drift; the center reads what is holding.

License

Ethotechnics Institute materials are published under CC BY-SA 4.0 . Credit the Institute, and publish anything you adapt under the same license. Entries carry citation metadata, so you can reference exactly what you used.

Studio

Evaluating a healthcare AI system? Ethotechnics Studio does commissioned safety evaluation for healthcare AI: a safeguards review, a readiness sprint on one workflow, investor diligence on a healthcare AI deal, or a clinical AI safety evaluation against FDA and EU AI Act expectations. That work runs through Ethotechnics Studio →

ethotechnics.org · open standards · scored public failures · free diagnostics