There is no checking step
Here is the uncomfortable part of the last lesson, stated plainly.
Producing a true sentence and producing a false one use exactly the same machinery. The model assembles likely-sounding continuations. Most of the time, likely-sounding and true line up, because the text it was trained on was mostly written by people trying to say true things. When they come apart, nothing notices, because there is nothing whose job is to notice.
This is usually called hallucination. That word is a bit flattering. The model is not malfunctioning when it invents a citation. It is doing precisely what it always does.
What it looks like in practice
In 2023 two lawyers in New York filed a court brief containing six cases that did not exist. A chatbot had produced them: plausible names, plausible courts, plausible years, plausible quotations. The lawyers were fined. The cases had the exact texture of real citations, because the model had absorbed the texture of thousands of real ones.
The same thing happens with a bus route that sounds right for a city it has read about, a customer helpline number with the correct number of digits, a section of a country's tax code with a convincing number, a research paper attributed to a real author who never wrote it.
Where it goes wrong most
There is a pattern, and it is useful.
Fabrication clusters on things that are specific and rare. Names, numbers, dates, page references, quotations, small local details, anything about a niche subject, anything about a person who is not famous.
It is much rarer on things that are general and repeated. Water boiling at 100 degrees at sea level appears in the training text in essentially that form, over and over, so it comes back reliably.
The reason follows from the mechanism. A pattern that appeared ten thousand times is deeply worn into the dials. A specific study's page number appeared once, or never, and the model still has to produce something, so it produces something of the right shape.
The cruel part: the confident tone is identical in both cases. Text about page numbers is written confidently by humans, so text about page numbers comes out confidently.
What does not help
Asking are you sure. You will get text about being sure. It has no separate access to its own reliability, so the question just prompts another round of likely-sounding continuation, often shaped by whether your tone suggested you wanted it to back down.
Asking for a source. If no search tool is attached, a request for a source is a request for text that looks like a source.
Assuming it improves with importance. It does not know which of your questions is the one that matters.
What does help
Attaching real search or real documents helps a great deal, and this is what most serious products now do. It does not end the problem, because the model can still misread or overstate what it retrieved, but you get links you can open, and opening them is the point.
Beyond that, a working rule: the more specific and checkable a claim is, the more you have to check it. Which is the opposite of the instinct. Specific claims feel more trustworthy. Here, specificity is the warning.
This is genuinely getting better and is not solved. Anyone who tells you the problem is fixed is selling something.
Before you move on