Casebook

Five public failures, scored

Each case was established by a court, an inquiry, or a regulator. Each is scored against the six state variables the laws track, with the clause that would have caught the drift, and against whether the failure trajectory ended in institutional learning or was absorbed as handled cases. The scores are not a verdict on anyone; they show where the coupling broke.

Six variables, five cases

Where the coupling broke

Read down a column to see which variable fails most often; read across a row to see the shape of one case. Standing has never held. Correction held once, and only because the halt was thrown before dependence set. The last column asks whether the failure trajectory ended in institutional learning or was absorbed as handled cases.

Case Capability Authority Evidence Dependency Standing Correction Time to halt Learned or handled
Robodebt Australia · July 2016 to November 2019 drifted failed failed drifted failed failed Three years and four months absorbed
The childcare benefits scandal Netherlands · Roughly 2012 to 2019 drifted drifted failed drifted failed failed About seven years partial
Post Office Horizon United Kingdom · 1999 to 2015 (prosecutions); redress continuing drifted drifted failed failed failed failed About twenty years absorbed
England's 2020 exam grades England · 13 to 17 August 2020 held drifted failed held failed drifted Four days absorbed
Apple Card credit limits United States · November 2019 to 2021 held held drifted held failed held About seventeen months to a policy change; the model was not withdrawn. learned
Failed, of 5 0 1 4 1 5 3

held explicit and coupled · drifted present but detached · failed absent, or what the inquiry found

Learned or handled — where the failure trajectory ended: learned = the exception changed the machinery · partial = some revision, late or imposed · absorbed = handled as cases, machinery unchanged. Read the theory

Empirical Record, drawn

State Variable Failure Rates Across 5 Statutory Inquiries

Public inquiries establish that automated systems rarely fail at technical capability. Failure concentrates where standing is denied, evidence is unbundled, and correction is withheld.

Held (Explicit & Coupled) Drifted (Detached) Failed (Absent / Broken) Capability 3 drifted , 2 held Authority 1 failed , 3 drifted , 1 held Evidence 4 failed , 1 drifted Dependency 1 failed , 2 drifted , 2 held Standing 5 failed Correction 3 failed , 1 drifted , 1 held 0% of cases 50% 100% (5 of 5 cases)

The cases

What the record established

Each case opens with the system in one line, the scale the inquiry found, and who finally made it stop.

  1. Robodebt

    Australia · Services Australia (Centrelink) · July 2016 to November 2019

    Automated welfare-debt raising: annual tax-office income averaged across fortnights to assert overpayments, with the burden of disproof placed on the recipient.

    Scale
    About 470,000 debts raised unlawfully. The class-action settlement the Federal Court approved in 2021 covered roughly 381,000 people, with repayments, wiped debts, and interest valued at about A$1.8 billion.
    Time to halt
    Three years and four months
    Halted by
    The Federal Court, on a consent order the Commonwealth agreed to hours before a hearing it would have lost.
    Learned or handled
    absorbed

    drifted Capability failed Authority failed Evidence drifted Dependency failed Standing failed Correction

    Read the scoring →

  2. The childcare benefits scandal

    Netherlands · Belastingdienst/Toeslagen · Roughly 2012 to 2019

    Risk-scored fraud detection on childcare benefit claims, with an all-or-nothing rule that reclaimed the entire benefit for any irregularity and an intent label that barred repayment arrangements.

    Scale
    More than 30,000 parents wrongly accused of fraud and made to repay benefits, often tens of thousands of euros; children placed in care; the cabinet resigned.
    Time to halt
    About seven years
    Halted by
    The Council of State reversing its own case law in October 2019, then a parliamentary inquiry.
    Learned or handled
    partial

    drifted Capability drifted Authority failed Evidence drifted Dependency failed Standing failed Correction

    Read the scoring →

  3. Post Office Horizon

    United Kingdom · Post Office Ltd, Fujitsu · 1999 to 2015 (prosecutions); redress continuing

    A branch accounting system whose reported shortfalls were treated as proof of theft or false accounting, by an operator that was also the investigator and the prosecutor.

    Scale
    More than 900 prosecutions; hundreds imprisoned, bankrupted, or both; the inquiry's first volume links at least thirteen suicides to the scandal and identifies roughly 10,000 eligible for redress.
    Time to halt
    About twenty years
    Halted by
    A group of 555 subpostmasters in civil litigation, then the Court of Appeal, then an Act of Parliament quashing convictions in bulk.
    Learned or handled
    absorbed

    drifted Capability drifted Authority failed Evidence failed Dependency failed Standing failed Correction

    Read the scoring →

  4. England's 2020 exam grades

    England · Ofqual, Department for Education · 13 to 17 August 2020

    A standardisation model that replaced cancelled A-level and GCSE exams by fitting each school's historical grade distribution to its current cohort, overriding teachers' assessed grades for all but the smallest classes.

    Scale
    About 39% of A-level grades issued below the teacher-assessed grade; the effect fell hardest on large cohorts in state schools and lightest on small classes, which were exempt.
    Time to halt
    Four days
    Halted by
    The Secretary of State, after Scotland had already reversed its equivalent and universities had begun allocating places on the model's grades.
    Learned or handled
    absorbed

    held Capability drifted Authority failed Evidence held Dependency failed Standing drifted Correction

    Read the scoring →

  5. Apple Card credit limits

    United States · New York Department of Financial Services · November 2019 to 2021

    Automated credit-limit decisions on a consumer card issued by Goldman Sachs, with no route by which an applicant could learn why their limit differed from a spouse's or ask for the decision to be reconsidered.

    Scale
    No unlawful discrimination found. The regulator's finding was that applicants, and the bank's own staff, could not explain individual outcomes, and that no reconsideration path existed.
    Time to halt
    About seventeen months to a policy change; the model was not withdrawn.
    Halted by
    Nobody. The issuer changed its policies after a regulator's investigation found the process, not the model, deficient.
    Learned or handled
    learned

    held Capability held Authority drifted Evidence held Dependency failed Standing held Correction

    Read the scoring →

Method

How to read a score

A score is a claim about one variable at the time of the harm, made from the primary record. It is not a finding of fault, and it is not a measurement of the system today.

held
The variable was explicit in the record and stayed coupled to the variables it should track. Capability was declared; authority named its ceiling; evidence was the kind that could justify the action taken.
drifted
The variable existed but detached from its neighbour. Authority that outlived the evidence it was issued on. A capability that grew twentyfold while the review capacity beside it stayed still.
failed
The variable was absent, or its absence is what the inquiry found. Standing that existed on paper and produced no state transition. Evidence that was never held by the party asserting it.

Every finding is pinned to the clauses whose obligation it falls under and to the law it instantiates. The clause links open the register on the standard's page. If a case is scored wrongly, the primary sources are listed on its page; tell us which one we misread.

The final verdict on each case is not a state variable. It is a claim about the trajectory: whether the sequence of failure, burden, challenge, and institutional response ended in case resolution or in institutional revision. A case can score badly on the variables and still end in learning, or score cleanly and still be absorbed. The distinction is exception absorption versus exception learning, and it is the subject of the theory essay.

For incidents written as remediation checklists rather than scored records, see the incident retrospectives.