The tells, and why they are disappearing
The tells, as they stood
There has been a folk practice of spotting generated images, and the list is worth writing down precisely — partly because some of it still works, and mostly to see how it has decayed.
- Hands. Six fingers, fused digits, thumbs on the wrong side. Largely fixed on current models; still fails in complex interactions such as two people holding hands.
- Text. Letter-like marks that spell nothing. Substantially fixed on large models, and still reliable on small local ones.
- Teeth and ears. Too many teeth, ears of different shapes. Mostly fixed.
- Backgrounds. Melted faces in crowds, structures that do not connect, patterns that shift. Still common, because this is the resolution limit from the earlier lesson rather than a training gap.
- Jewellery and eyewear. Asymmetric earrings, spectacle arms that vanish. Still frequent, because it is the symmetry-and-continuity failure.
- The surface. Skin without pores, uniform lighting, a general glossiness. Still the strongest single tell, and it is the distribution pull rather than any defect.
Notice how the list divides. The failures that were about training coverage got fixed by more data. The failures that are structural — global consistency, latent resolution, mode-seeking — have not been fixed, because more data does not address them.
Why "I can always tell" is not true
Two facts make confident visual identification unreliable, and both are uncomfortable.
Selection. You see the generated images that were noticed. Every convincing one passed by without registering. This is survivorship bias operating on your entire sense of how good these systems are, and there is no way to correct for it by looking harder.
The failures that remain are removable. A skilled retoucher fixes hands, adds grain, breaks the lighting, and photographs a screen to add real optical noise. Every tell on the list above is a fifteen-minute job for someone who intends to deceive. The tells identify casual generation, which is a different and much easier problem than identifying deliberate deception.
The consequence is that a false accusation is now a real risk. Photographers have had genuine work called AI-generated on the basis of smooth skin, which is what studio lighting and ordinary retouching produce. The costs land on individuals with no route to prove a negative.
What to use instead
The next-to-last module covers provenance properly. The short version, which is worth having now:
Verify the source and the claim, not the pixels. Where did this file come from, who published it first, does anyone with a name stand behind it, does a reverse image search show it appearing earlier somewhere else. Free tools — Google Lens, TinEye, Yandex — answer the last question in seconds and answer it far more reliably than any inspection of fingers.
Treat an image with no traceable origin as unverified, whether or not it looks generated. That is the same standard journalism used before any of this existed, and it survives every improvement in the models.
The habit worth building
When you catch yourself thinking "that looks AI", finish the thought properly: which mechanism would have produced that artefact? Melted background faces mean latent resolution. Wrong reflections mean no global constraint. Glossy uniformity means distribution pull and high guidance.
If you can name the mechanism, you have a real observation. If you cannot, you have a feeling — and feelings about images are exactly what everybody's is being trained on at the moment, by both the generators and the people who benefit from doubt.
What to say when someone asks
You will be asked, because everybody with any reputation for knowing about this gets asked. A good answer has three parts and takes fifteen seconds.
Say what you can see and what mechanism it suggests, if anything. Say that visual inspection cannot establish origin either way, and that people have been wrongly accused on exactly this basis. Then move to the question that can be answered: where did the file come from, and who is standing behind it.
That answer is unsatisfying to somebody who wanted a verdict, and it is the truthful one. Declining to give a confident verdict on an image is not a failure of expertise. At the moment it is what expertise consists of.
The one thing to keep
The visible artefacts people use to identify generated images are transient consequences of specific architectures, so a detection habit built on them expires and confident identification by eye is not reliable.
Before you move on
Why has the "count the fingers" heuristic become unreliable, while melted faces in crowd backgrounds remain common?
Pick the one you would defend. Nobody sees your answer.