Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

AI, Safety and What Goes Wrong

The failure modes of AI, stated plainly, with the numbers.

Lesson 4 of 738 min

What a benchmark score does not tell you

The bar exam problem

When a model is announced, it arrives with numbers: a percentile on a professional exam, a percentage on a reasoning set, a bar on a chart taller than last year's bar. These numbers are usually real. They are also, for the question you care about, close to useless on their own — and knowing why is one of the most transferable skills in this course.

Start with the famous claim that a model passed a bar exam near the top of the range. Two things were true at once. The model did score well on multiple-choice and essay material. And the percentile was computed against a pool that included repeat takers on a February sitting, which is not the pool people picture when they hear "top ten per cent". A later re-analysis put the figure far lower against first-time July takers. Nobody fabricated anything. The comparison group was chosen, and the choice moved the headline by tens of percentiles.

That is the general shape. A benchmark result is a measurement of a specific thing, under specific conditions, against a specific comparison — and the summary sentence drops all three.

Four reasons the number does not transfer

Contamination. Benchmarks are published on the internet. Training sets are scraped from the internet. If the test questions and answers are in the training data, the model is being examined on material it has already seen. Labs try to remove known benchmarks, but paraphrases, forum discussions and solution write-ups survive. When a model does markedly better on a benchmark released before its cutoff than on an equivalent one released after, contamination is the first explanation to reach for — and researchers have repeatedly found exactly that pattern.

The task is not your task. Exams are designed to be gradeable: self-contained questions, one right answer, no missing information, no consequences for a wrong one. Your work is the opposite. The information is incomplete, the question is badly posed by the person who asked it, the right answer depends on context nobody wrote down, and being wrong costs something. A model can be excellent at the first and poor at the second.

The conditions were generous. Reported numbers often come from the best of several attempts, or with a carefully tuned prompt, or with an expensive setting you are not using, or with tool access. Multiple sampled attempts scored as "correct if any is correct" is a legitimate research measure and a misleading product claim. Check whether the number was one shot.

One model, four ways to report the same benchmarkBest of ten attempts,tuned prompt, tools on88One attempt, tuned prompt79One attempt, the promptyou would type71Thirty examples from yourown work58% correctIllustrative. The headline sentence keeps only the first bar. The number that predicts your outcome isthe last one, and it is the only one nobody publishes for you.
One model, four ways to report the samebenchmarkBest of ten attempts, tuned prompt, tools on88One attempt, tuned prompt79One attempt, the prompt you would type71Thirty examples from your own work58% correctIllustrative. The headline sentence keeps only thefirst bar. The number that predicts your outcome isthe last one, and it is the only one nobodypublishes for you.

Saturation. When everyone optimises against the same test, the test stops discriminating. Scores bunch at the top and the remaining differences are noise, or worse, are differences in how well each lab gamed the format. This is Goodhart's law with a leaderboard: a measure that becomes a target stops being a good measure.

What to ask instead

You do not need to be a researcher to interrogate a claim. Four questions cover most of it.

  1. Measured on what, exactly? Name the benchmark and go and read three of its questions. This takes four minutes and is startlingly clarifying — many famous benchmarks contain items that are ambiguous or mislabelled.
  2. Compared with whom, and when? Against another model on the same date, or against a human population that may not be the population implied?
  3. How many attempts, and with what scaffolding?
  4. Would a difference of this size change my decision? Two points on a benchmark almost never should.

The test that actually matters

Build your own. Twenty to fifty real examples from your own work, with answers you have checked yourself. Run each candidate model over them. This is not a scientific evaluation — it is too small for that, and the next course in this track is about doing it properly — but it beats every public leaderboard for your specific decision, because it measures the thing you will actually do.

Most people are surprised twice. First that the model ranked lower on the public chart does better on their examples. Second that their own examples were harder to write down than expected, because much of what makes an answer good at their job had never been stated out loud.

The one thing to keep

A benchmark measures a specific task under generous conditions against a chosen comparison group, and twenty examples from your own work predict your outcome better than any public leaderboard.

Before you move on

Model A scores 91% on a widely used reasoning benchmark; Model B scores 88%. On your team's fifty real support tickets, B is clearly better. What is the most likely explanation?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly