When the state automates suspicion
Two national failures
The most instructive automated-decision disasters so far both involved governments accusing citizens of fraud, at scale, wrongly.
The Netherlands, childcare benefits. From around 2013 the Dutch tax authority used risk models to flag applications for childcare allowance as potentially fraudulent. Dual nationality and low income were among the factors that raised risk. Flagged families had their benefits stopped and were ordered to repay years of allowance in full, often tens of thousands of euros, frequently for administrative errors such as a missing signature. Appeals were slow and the accusation carried a fraud designation that closed off other support. Roughly 26,000 families were affected. Families were driven into debt and separation, and more than a thousand children were placed in foster care. In January 2021 the entire Dutch government resigned over it. The Dutch data protection authority later fined the tax administration for unlawful processing, including the discriminatory use of nationality.
Australia, Robodebt. From 2015 the Australian government matched welfare recipients' reported fortnightly income against annual tax data, and where the annual figure was higher, averaged it across the year to infer undeclared income and raise a debt. The averaging assumption is simply wrong for anyone with irregular work — precisely the population receiving these payments. Hundreds of thousands of debts were raised, the burden of disproof was placed on recipients, and debt collectors were used. A 2023 Royal Commission called the scheme unlawful and described it in scathing terms; more than a billion Australian dollars was refunded and a large settlement paid.
The pattern, which is not about the algorithm
Neither failure required sophisticated technology. Robodebt was arithmetic. What made them catastrophic was a set of design decisions that recur:
The burden of proof was inverted. The system's output was treated as an established debt, and the citizen had to disprove it — often using payslips from years earlier that nobody keeps.
The error was systematic, not random. A flawed assumption applied identically to everyone it touched, so the mistakes were correlated and concentrated on the people least able to resist.
Appeal was slower than harm. Money was withheld or collected during the appeal. A process that takes fourteen months to vindicate you has already done the damage.
Warnings existed and were routed around. Internal legal advice in the Australian case had questioned lawfulness. In the Dutch case, journalists and parliamentarians raised individual cases for years.
Nobody owned the decision. Responsibility was distributed across the model, the policy, the operational rules and the caseworker, so every individual could truthfully say the decision was not theirs.
The legal counterweight
One case is worth knowing by name. In February 2020 the District Court of The Hague ruled that SyRI, a Dutch system that combined data across government departments to generate fraud risk scores in specific neighbourhoods, violated Article 8 of the European Convention on Human Rights. The court found the legislation failed the test of striking a fair balance: it was insufficiently transparent and verifiable, and it targeted poorer areas, risking discrimination and stigmatisation.
That judgment is the clearest statement available that opacity itself can be unlawful — not because the system was proven inaccurate, but because citizens could not know how they were assessed. The UN Special Rapporteur on extreme poverty intervened in the case and has written about the emergence of a "digital welfare state" in which the poor are surveilled as a condition of receiving support.
What follows for anyone building or buying this
Five requirements, each of which was missing above.
Never let a flag be a finding. A risk score is a reason to look, and looking must involve a human with the file and the time.
Do not withhold the entitlement during investigation unless there is specific individual evidence. The default should protect the person, not the budget.
Test the assumption against the population it will hit. Income averaging fails for irregular earners; anyone who had run the arithmetic on a sample of real cases would have seen it in a day.
Publish the criteria. If a factor cannot be defended in public, it will not survive a court either.
Count the false accusations, and publish that too. Every one of these programmes reported recoveries. None reported how many innocent people it accused, because nobody was required to measure it, and what is not measured does not exist until a Royal Commission creates it.
The one thing to keep
The catastrophic welfare systems failed not through clever algorithms but through inverted burden of proof, correlated errors, appeals slower than the harm and diffused ownership — and a Dutch court found that opacity alone made such a system unlawful.
Before you move on
Robodebt inferred debts by averaging annual tax income across fortnights. Why was this so damaging specifically for the population it was applied to?
Pick the one you would defend. Nobody sees your answer.