Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Making Things With AI

Images, video, voice and music — how they work, where they break, who owns them.

Lesson 23 of 848 min

The pull toward the average

Everyone has seen this and few people name it

Generate a photo of a woman twenty times on an unmodified model and look at the set. The faces will differ, and they will differ around a centre: a similar age, a similar symmetry, similar skin, similar lighting, a similar pleasant expression. The same happens with a house, a meal, an office.

This is the distribution pull, and it has two separate causes that compound.

The prompt underspecifies. Everything you do not say is filled in by what is most probable. Most probable is not neutral. It is the mode of a corpus that heavily over-represents commercial photography, stock imagery and social media, all of which are already selected for a narrow idea of what looks good.

Guidance amplifies the typical. Recall the arithmetic: guidance pushes away from the unconditional prediction and toward the conditional one, past where either pointed. That extrapolation drives every region toward the most prototypical version of what it is. High guidance is, quite literally, a dial for how stereotypical the output is.

The visible signature

Once you know what to look for, the pull is easy to spot:

  • Skin without pores, blemishes or asymmetry.
  • Teeth uniformly white and even.
  • Lighting that is always flattering and always from a plausible key with fill.
  • Objects that are new, clean and undamaged.
  • Food that is styled.
  • Interiors that are tidy.

None of this is a rendering defect. Every element is a plausible thing. It is the combination that never occurs in an unposed photograph, and it is why generated images read as advertising even when the subject is mundane.

Working against it

Three approaches, in increasing order of effectiveness.

Specify the atypical. Say the things you would not have thought to say: the time of day, the state of the objects, the imperfection. A cluttered kitchen at 7am, unwashed pan on the hob, low winter light through a dirty window, someone's coat over a chair. Every clause you add takes a decision away from the mode.

Lower the guidance. Dropping from 8 to 4 does more for realism than most prompt vocabulary. It costs you prompt obedience, which is a real trade and worth making when the subject is simple.

Condition on a real photograph. A depth or edge map taken from an ordinary snapshot imports the composition of a real, unposed scene — which is the very thing the model will not produce on its own.

Why this matters more than an aesthetic complaint

The pull toward the average is the mechanism behind the bias covered in the next lesson, and it is worth seeing them as one thing. When a model produces the same kind of face for a doctor every time, it is not applying a rule about doctors. It is doing what it does for kitchens and lamps: returning the centre of a distribution. The distribution has a demographic shape because the corpus does.

Understanding it as one mechanism has a practical benefit: the same interventions work. Specification, lower guidance and reference conditioning all reduce demographic collapse for the same reason they reduce aesthetic collapse.

There is a limitation nobody has removed. You can push away from the mode, but you cannot get to a region the training data barely covers. Ask for a kind of face, a kind of room or a kind of tool that is rare in the corpus and specification will not conjure it; you will get the nearest well-covered thing with your adjectives applied on top. That boundary — between "underspecified" and "not there" — is the single most useful thing to be able to tell apart, and the test is simple: add the specification and see whether the image moves toward what you meant or merely acquires a label of it.

It is worth noticing that the pull is also why generated images age so visibly. The mode of the corpus at the time of training becomes the look of everything the model makes, so a model trained in one year produces a recognisable house style that dates the way a decade's stock photography dates. Anyone building a body of work on a single checkpoint is committing to that look. That is an argument for a light touch — generate elements, finish them yourself, keep your own grade — rather than for shipping the model's default and calling it a style.

The one thing to keep

Guidance and short prompts both push output toward the most typical example in the training distribution, which is why unspecified images converge on a narrow, glossy sameness.

Before you move on

Twenty generations of "a photo of a kitchen" all return tidy, evenly lit, magazine-like rooms. Which single change most directly counteracts the mechanism?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly