Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

AI, Safety and What Goes Wrong

The failure modes of AI, stated plainly, with the numbers.

Lesson 69 of 738 min

What an audit can and cannot see

The frameworks people actually use

Below the law sits a layer of voluntary standards, and these are what an organisation practically works from.

The NIST AI Risk Management Framework, published in 2023 with a generative AI profile added in 2024, is the most widely adopted. It is organised around four functions: Govern (policies, roles, accountability), Map (context and what could go wrong), Measure (analysis and tracking), Manage (prioritisation and response). It is voluntary, free, deliberately non-prescriptive, and useful mainly as a structure that stops teams from forgetting a category.

ISO/IEC 42001, published in 2023, is an AI management system standard, and unlike NIST's framework it is certifiable — an accredited body audits you and issues a certificate. It concerns whether you have a working management system, not whether any particular model is good.

ISO/IEC 23894 covers AI risk management guidance, and sits alongside.

The distinction matters and is often blurred in marketing. A certificate against 42001 says an organisation has processes. It does not say a model is accurate, fair or safe.

Documentation artefacts worth knowing

Three formats have become the vocabulary of transparency.

Model cards (proposed by Mitchell and colleagues in 2019) — a short document stating intended use, out-of-scope uses, training data at a high level, evaluation results disaggregated by group, and known limitations. The disaggregation is the important part: overall accuracy hides exactly what the second module taught you to look for.

Datasheets for datasets (Gebru and colleagues) — the equivalent for data: why it was collected, by whom, from where, with what consent, with what known gaps.

System cards — the same idea applied to a deployed product rather than a model, covering the whole pipeline including filters and human review.

These are genuinely useful and they are self-reported. Reading one tells you what the organisation chose to disclose, in a format that makes omissions visible to someone who knows what should be there.

What an audit actually is

Here is where realism is needed.

An audit sees what the auditee provides. Unless there is a legal power of inspection, the auditor receives selected documentation, selected data and selected access. Most AI audits are conducted with the client's cooperation on the client's material.

Scope is negotiated. A favourable audit is often a narrow one. Read the scope statement before the conclusions; it is where the interesting information is.

Standards for what is being tested barely exist. Financial audit rests on decades of accounting standards. There is no equivalent settled body for AI, so two auditors can reach different conclusions honestly.

The incentives run the wrong way. The auditee pays. This problem is well documented in financial and social auditing and there is no reason to expect this field to escape it.

Researchers call the failure mode audit washing: a certificate becomes evidence of responsibility and reduces external scrutiny, without changing the system.

New York City's Local Law 144 is the most instructive case. It requires an annual independent bias audit of automated employment decision tools with published impact ratios. Early research found low compliance and considerable variation in how audits were conducted and what they covered — which tells you both that mandating audits changes something and that mandating them is not sufficient.

How to read an audit or a card

Six questions.

  1. What was in scope, and what was excluded?
  2. Did the auditor have access to the model, the data, and production logs — or to documents about them?
  3. Are results disaggregated by group, or reported as a single figure?
  4. Are limitations named specifically, or in general language?
  5. Who paid, and does the auditor sell remediation services?
  6. What would a failing result have looked like? If no result would have failed, it was not a test.

That last question is the one that separates assurance from theatre, and it applies well beyond AI.

What is worth doing anyway

Despite all of the above, the documentation habit is worth adopting, for a reason that has nothing to do with certificates: writing down intended use, out-of-scope use, evaluation by group and known limitations forces a team to discover what it does not know. Most of the value is produced during the writing, by the people writing it, before any auditor arrives.

The one thing to keep

A management-system certificate says an organisation has processes, not that a model is safe — read an audit's scope and access first, and ask what a failing result would have looked like.

Before you move on

A vendor presents an ISO/IEC 42001 certificate as evidence that its hiring model is fair. What does the certificate actually establish?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly