Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Hallucination, and how to check

AI, Safety and What Goes Wrong · lesson 3 of 8 · 9 min

A language model predicts likely next words. That is the whole mechanism. It is astonishingly powerful and it explains the central failure: the model will produce a fluent, well-formed, completely fabricated answer, and it will sound exactly like the true ones.

The word "hallucination" is a bit misleading. Nothing unusual is happening inside. The model is doing the same thing when it is right and when it is wrong.

What it looks like

In 2023, two New York lawyers submitted a court filing citing six cases. None existed. The model had produced case names, docket numbers, quotes and internal citations, all in the correct format. The judge fined them $5,000. The filing did not look sloppy. It looked professional, which was the trap.

Fabrication clusters in predictable places:

  • Citations, references and URLs. The format is highly patterned, so it is easy to generate and hard to eyeball as fake.
  • Numbers and dates. Population figures, drug dosages, exchange rates, historical years.
  • Names attached to facts. Who founded what, who said which quote, who wrote which paper.
  • Anything niche. The less a topic appeared in training data, the more the model is improvising. A question about a small town's municipal rules is far riskier than one about photosynthesis.
  • Anything recent. Models have a training cutoff. Ask about last month and you may get a confident answer built from the shape of older events.

The confidence trap

Asking "are you sure?" does very little. The model will often apologise and change its answer whether or not the original was correct — it is predicting what a helpful response to doubt looks like, not re-examining evidence. Some systems now show confidence scores or hedging language, and these are weakly informative at best.

Treat stated confidence as a writing style, not as evidence.

How to actually check

A working method, in order of effort:

Ask for the source, then open it. Not "give me a source" — that invites a plausible-looking fabricated one. Copy the exact title into a search engine and see whether it exists and whether it says what was claimed. A real paper that says something different is as much a failure as a fake one.

Ask the same question twice in separate conversations. Facts the model actually holds tend to come back stable. Fabrications drift — different year, different name. This is a cheap smoke test, not proof.

Give it the document instead of asking from memory. Paste the annual report and ask questions about the text in front of it. Grounding the model in supplied material sharply reduces invention, though it does not eliminate it — models still misread and over-summarise. (Consider the privacy cost first; that is the next lesson.)

Check the claim, not the paragraph. Fluent text carries wrong facts inside right ones. Pull out each checkable assertion separately.

The rule worth keeping

Match your verification to the consequence of being wrong. Brainstorming names for a shop: no verification needed. Drafting an email: skim it. A dosage, a legal deadline, a figure going into a report your manager will sign, a claim you will state publicly: verify every specific, from a source you opened yourself.

The useful mental model is not "an assistant that sometimes errs". It is "a very well-read colleague who never says *I don't know*". You would still check their citations.

Before you move on

You ask a model for statistics on youth unemployment in Kenya. It gives a precise figure, names a report, and adds "I am confident in this figure." You ask "are you certain?" and it confirms. What have you learned?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly