Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

AI, Safety and What Goes Wrong

The failure modes of AI, stated plainly, with the numbers.

Lesson 36 of 7310 min

When the state automates suspicion

Two national failures

The most instructive automated-decision disasters so far both involved governments accusing citizens of fraud, at scale, wrongly.

The Netherlands, childcare benefits. From around 2013 the Dutch tax authority used risk models to flag applications for childcare allowance as potentially fraudulent. Dual nationality and low income were among the factors that raised risk. Flagged families had their benefits stopped and were ordered to repay years of allowance in full, often tens of thousands of euros, frequently for administrative errors such as a missing signature. Appeals were slow and the accusation carried a fraud designation that closed off other support. Roughly 26,000 families were affected. Families were driven into debt and separation, and more than a thousand children were placed in foster care. In January 2021 the entire Dutch government resigned over it. The Dutch data protection authority later fined the tax administration for unlawful processing, including the discriminatory use of nationality.

Australia, Robodebt. From 2015 the Australian government matched welfare recipients' reported fortnightly income against annual tax data, and where the annual figure was higher, averaged it across the year to infer undeclared income and raise a debt. The averaging assumption is simply wrong for anyone with irregular work — precisely the population receiving these payments. Hundreds of thousands of debts were raised, the burden of disproof was placed on recipients, and debt collectors were used. A 2023 Royal Commission called the scheme unlawful and described it in scathing terms; more than a billion Australian dollars was refunded and a large settlement paid.

Two states automate suspicion2013The Dutch tax authority starts flagging childcare benefit claims with risk models;dual nationality raises the score2015Australia's Robodebt averages annual tax income across fortnights and raises debtsagainst welfare recipientsFeb 2020The Hague court rules SyRI unlawful: too opaque to verify, aimed at poorerneighbourhoodsJan 2021The Dutch government resigns; about 26,000 families affected, more than a thousandchildren placed in care2023A Royal Commission finds Robodebt unlawful; more than a billion Australian dollarsrefundedNeither failure needed clever technology; Robodebt was arithmetic. What made them catastrophic was aninverted burden of proof, correlated errors, appeals slower than the harm, and nobody who owned thedecision.
Two states automate suspicion2013The Dutch tax authority starts flaggingchildcare benefit claims with risk models; dualnationality raises the score2015Australia's Robodebt averages annual tax incomeacross fortnights and raises debts againstwelfare recipientsFeb 2020The Hague court rules SyRI unlawful: too opaqueto verify, aimed at poorer neighbourhoodsJan 2021The Dutch government resigns; about 26,000families affected, more than a thousandchildren placed in care2023A Royal Commission finds Robodebt unlawful;more than a billion Australian dollars refundedNeither failure needed clever technology; Robodebtwas arithmetic. What made them catastrophic was aninverted burden of proof, correlated errors, appealsslower than the harm, and nobody who owned thedecision.

The pattern, which is not about the algorithm

Neither failure required sophisticated technology. Robodebt was arithmetic. What made them catastrophic was a set of design decisions that recur:

The burden of proof was inverted. The system's output was treated as an established debt, and the citizen had to disprove it — often using payslips from years earlier that nobody keeps.

The error was systematic, not random. A flawed assumption applied identically to everyone it touched, so the mistakes were correlated and concentrated on the people least able to resist.

Appeal was slower than harm. Money was withheld or collected during the appeal. A process that takes fourteen months to vindicate you has already done the damage.

Warnings existed and were routed around. Internal legal advice in the Australian case had questioned lawfulness. In the Dutch case, journalists and parliamentarians raised individual cases for years.

Nobody owned the decision. Responsibility was distributed across the model, the policy, the operational rules and the caseworker, so every individual could truthfully say the decision was not theirs.

The legal counterweight

One case is worth knowing by name. In February 2020 the District Court of The Hague ruled that SyRI, a Dutch system that combined data across government departments to generate fraud risk scores in specific neighbourhoods, violated Article 8 of the European Convention on Human Rights. The court found the legislation failed the test of striking a fair balance: it was insufficiently transparent and verifiable, and it targeted poorer areas, risking discrimination and stigmatisation.

That judgment is the clearest statement available that opacity itself can be unlawful — not because the system was proven inaccurate, but because citizens could not know how they were assessed. The UN Special Rapporteur on extreme poverty intervened in the case and has written about the emergence of a "digital welfare state" in which the poor are surveilled as a condition of receiving support.

What follows for anyone building or buying this

Five requirements, each of which was missing above.

Never let a flag be a finding. A risk score is a reason to look, and looking must involve a human with the file and the time.

Do not withhold the entitlement during investigation unless there is specific individual evidence. The default should protect the person, not the budget.

Test the assumption against the population it will hit. Income averaging fails for irregular earners; anyone who had run the arithmetic on a sample of real cases would have seen it in a day.

Publish the criteria. If a factor cannot be defended in public, it will not survive a court either.

Count the false accusations, and publish that too. Every one of these programmes reported recoveries. None reported how many innocent people it accused, because nobody was required to measure it, and what is not measured does not exist until a Royal Commission creates it.

The one thing to keep

The catastrophic welfare systems failed not through clever algorithms but through inverted burden of proof, correlated errors, appeals slower than the harm and diffused ownership — and a Dutch court found that opacity alone made such a system unlawful.

Before you move on

Robodebt inferred debts by averaging annual tax income across fortnights. Why was this so damaging specifically for the population it was applied to?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly

When the state automates suspicion · AI, Safety and What Goes Wrong · Addaly