Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

AI, Safety and What Goes Wrong

The failure modes of AI, stated plainly, with the numbers.

Lesson 23 of 738 min

How professional checkers actually work

The counter-intuitive finding

In a study at Stanford, historians, undergraduates and professional fact-checkers were given unfamiliar websites and asked to judge their reliability. The historians and students did what educated people do: they read the site carefully, examined its design, looked at its About page, weighed its arguments. They were frequently fooled.

The fact-checkers were fast and accurate, and the reason was that they barely read the site at all. Within seconds they had left it, opened new tabs, and were reading what other sources said about the organisation. The researchers called this lateral reading, as opposed to reading vertically down the page.

The lesson generalises directly to AI output. Scrutinising a fluent paragraph for signs of falsehood is the historians' strategy, and it fails for the same reason: a well-constructed page and a well-constructed paragraph both look exactly like a good one. Leave, and check elsewhere.

Reading down the page against reading sidewaysVertical reading: the historians and undergraduatesRead the site carefully, top to bottomExamined the design and the About pageWeighed the arguments on their meritsFrequently fooledLateral reading: the professional fact-checkersLeft the page within secondsOpened new tabs on the organisation and theclaimRead what independent sources said about itFast, and accurateThe historians and students were fooled; the fact-checkers were fast and right. A fluent paragraph anda well-built page both look exactly like good ones, which is why judging them by reading harder doesnot work.
Reading down the page against readingsidewaysVertical reading: the historians andundergraduatesRead the site carefully, top to bottomExamined the design and the About pageWeighed the arguments on their meritsFrequently fooledLateral reading: the professionalfact-checkersLeft the page within secondsOpened new tabs on the organisation and theclaimRead what independent sources said about itFast, and accurateThe historians and students were fooled; thefact-checkers were fast and right. A fluentparagraph and a well-built page both look exactlylike good ones, which is why judging them by readingharder does not work.

The core move, in four steps

Pull out the checkable atom. A paragraph is not checkable; a claim is. "The scheme covered 1.2 million households by March 2023" is an atom. Extract them one at a time, because fluent text glues true and false atoms together and reading for overall impression cannot separate them.

Go sideways, not down. Search the specific claim, in its own words, plus a term that would bring up the original: the scheme's name, the ministry, the journal. Do not search for confirmation of the sentence — search for the entity and read what it says.

Reach the primary document. Court filing, statute, annual report, dataset, press release, published paper. This is one or two clicks in most cases and it settles most disputes immediately.

Check what the claim omits. Correct-but-misleading is more common than false. A real figure from a real report describing a different population, year or definition.

Tools, all free

None of this requires a subscription.

  • Reverse image search — Google Lens, TinEye, Bing Visual Search, Yandex. For any photograph presented as evidence, this is the first move. An image that has been online for six years is not from yesterday's event.
  • The Wayback Machine — what a page said before it was edited, and whether it existed at all. Also the answer when a source has been quietly deleted.
  • Frame-level video checking — take a screenshot of a distinctive frame and reverse-search that. Repurposed old footage is by volume a far more common problem than deepfakes.
  • Site registration data — a domain registered three weeks ago and presenting itself as a decades-old institute is settled with one lookup.
  • Statistical offices and official registries — India's data portal, Eurostat, the ONS, the World Bank, and national company registries. They publish the tables underneath the headlines and they are free.
  • Semantic Scholar, PubMed, arXiv and Google Scholar — for whether a paper exists at all, and for its abstract, which contains the population and the effect size.

Where this fits with AI use

A workable routine for a claim that matters:

First, ask what kind of claim is this? Recent, niche, numerical, or about a named person — those are the four high-risk categories from the hallucination lesson, and they deserve checking before anything else.

Second, extract the atoms and check each with a lateral search.

Third, if two independent sources disagree, do not average them. Find out why they disagree, because the reason is usually a definition — different years, different inclusion criteria, different geography — and understanding it is more valuable than either number.

And fourth, know when to stop. Verification cost should track consequence. Two minutes for something you will say in a meeting; twenty for something going into a document with your name on it; more if someone could be harmed. The triage from the first module is exactly the schedule for this.

The habit worth building

Speed comes from doing it in a fixed order rather than from working harder. Atom, sideways search, primary source, omission check. Four steps, most of the time under three minutes.

The people who are good at this are not more sceptical than you. They have simply stopped trying to judge a text by reading it more carefully, which is the move that feels responsible and does not work.

The one thing to keep

Judging a claim by reading it closely fails; professionals leave the page immediately, extract one checkable claim at a time, and read what independent primary sources say about it.

Before you move on

You are sent a striking photograph said to show flooding in a named city yesterday. What is the highest-value first check?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly