Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

AI, Safety and What Goes Wrong

The failure modes of AI, stated plainly, with the numbers.

Lesson 70 of 738 min

Learning from what went wrong

The comparison that keeps being made

Commercial aviation is extraordinarily safe, and the usual explanation is technology. The better explanation is reporting. Aviation built, over decades, a system in which failures are reported, investigated by an independent body, published in full, and turned into changes that everyone must adopt.

Two features do the work.

Mandatory investigation with independence. Accident investigators are separate from the regulator and from the operators, publish findings including uncomfortable ones, and their remit is causes rather than blame.

Confidential, non-punitive voluntary reporting. The US Aviation Safety Reporting System, run by NASA rather than the aviation regulator specifically so that reporting does not go to the enforcer, receives many tens of thousands of reports a year from pilots and controllers describing their own mistakes. Reporting confers limited protection from enforcement action. That protection is why the reports exist, and the near-misses are where most of the learning is.

AI has neither at present. What it has is journalism, litigation and researchers, which is a poor substitute because all three are adversarial, and adversarial channels select for the cases somebody can be blamed for rather than the cases everybody could learn from.

What does exist

The AI Incident Database, run by the Responsible AI Collaborative, catalogues publicly reported incidents — over a thousand and growing — with structured fields, and is free to search. It is compiled from public sources, so it captures what was reported and not what happened.

The OECD's AI Incidents Monitor tracks incidents from news sources internationally, and the OECD has been working on a common incident reporting framework so that different countries' regimes can share definitions.

The EU AI Act creates a duty. Providers of high-risk systems must report serious incidents to authorities. This is the first significant mandatory regime and it applies to a defined subset, not to everything.

Sectoral regimes already bite. Medical device regulators have adverse-event reporting that covers AI-based devices. Financial regulators require operational incident reporting. Data protection authorities require breach notification, often within 72 hours.

What counts as an incident

Broader than most people assume, and defining it narrowly is how organisations end up reporting nothing.

A harm that occurred. A near-miss caught before it reached anyone. A system behaving outside its stated scope. A material accuracy degradation. A discovered disparity between groups. An unauthorised use. A confidentiality breach through an AI tool. A supplier's model changing under you.

The near-misses are the valuable ones, because they are frequent and cheap. Any reporting culture that only captures realised harm is capturing the small tail of its available information.

Building this where you are

You do not need a national framework to do this internally.

Define an incident in writing, including near-misses, and give examples so people recognise one.

Make reporting non-punitive and mean it. This is the whole thing. If the first report leads to a disciplinary process, it will be the last report, and thereafter you will have a clean record and no information. Aviation learned this the expensive way.

Record a fixed set of fields: what happened, which system and version, who was affected, how it was detected, what was done, what changed as a result. The last field is the one that makes people keep filing.

Convert incidents into tests. Every incident becomes a permanent case in your evaluation set. This is the single practice that stops the same failure recurring, and it is standard in software engineering and rare in AI deployment.

Review them together, on a schedule. Individually they are anecdotes; twenty of them are a pattern, and the pattern is usually not what anyone predicted.

Watch the count. Zero reported incidents means the reporting is broken, not that nothing happened. This is the third time this course has made that point, about override rates, about complaint counts and now here, because it is the most reliable way to be comfortably wrong.

A reporting loop you can build without a regulatorDefine anincident,includingnear-missesWith examplespeoplerecogniseMakereportingnon-punitive,and mean itThe firstdisciplinaryprocess endsthe reportsRecord fixedfieldsWhat, whichsystem andversion, whowas affected,how detected,what changedConvert eachincidentinto a testcaseA permanentitem in theevaluationsetReview themtogether, ona scheduleTwentyanecdotes area patternWatch thecountZero is abrokenchannelAviation's record rests on reporting that does not go to the enforcer. The near-misses are where thelearning is, and a count of zero means the reporting is broken, not that nothing happened.
A reporting loop you can build without aregulatorDefine an incident, including near-missesWith examples people recogniseMake reporting non-punitive, and mean itThe first disciplinary process ends the reportsRecord fixed fieldsWhat, which system and version, who wasaffected, how detected, what changedConvert each incident into a test caseA permanent item in the evaluation setReview them together, on a scheduleTwenty anecdotes are a patternWatch the countZero is a broken channelAviation's record rests on reporting that does notgo to the enforcer. The near-misses are where thelearning is, and a count of zero means the reportingis broken, not that nothing happened.

Reporting outward

When an incident affects people outside your organisation, the duty may be legal — data protection authorities, sectoral regulators, and under the EU regime, market surveillance authorities. Beyond duty, submitting to a public database is a small act with disproportionate value, because the field's collective knowledge of how these systems fail is currently assembled from whatever happened to become news.

The one thing to keep

Aviation's safety record rests on independent investigation and confidential non-punitive reporting of near-misses, and the equivalent for AI barely exists — so define incidents to include near-misses, make reporting safe, and turn every one into a permanent test case.

Before you move on

Why is the US aviation voluntary reporting system deliberately run by NASA rather than by the aviation regulator?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly