Nobody good ships a first generation
The workflow people actually use looks like this: generate until you have a base with the right composition and light, fix regions one at a time, extend the frame if you need to, upscale last. The prompt got you sixty percent. The rest is masking.
Inpainting, and the setting that changes everything
Inpainting means painting a mask over a region, describing what should be there, and regenerating only that region while the model reads the surrounding pixels for context.
The setting that separates a clean result from a smear is usually called "inpaint at full resolution" or "inpaint only masked area". Without it, the entire image is squeezed down to the model's working size, so a face occupying five percent of a wide shot gets maybe forty pixels and comes back as mush. With it, the tool crops around your mask, generates that crop at the model's native resolution, and blends it back. Same model, same prompt, completely different outcome.
Three practical rules:
- Mask a little wider than the problem. Tight masks leave visible seams.
- Change the prompt to describe the region, not the scene. If you are fixing a hand, prompt the hand.
- Fix one thing per pass, and save between passes.
Denoising strength is the dial you actually turn
Whether you are inpainting or running image-to-image, one number governs how far the result may travel from the input. It sets how much noise is added before denoising begins.
- 0.2 to 0.35 — texture and finish only. Shapes survive intact.
- 0.4 to 0.6 — the same objects with different details. Faces will change. Most real work happens here.
- 0.7 and up — the input is a loose suggestion. Composition survives, content does not.
If your edit is not taking, raise it. If your subject keeps turning into a different person, lower it and do two passes.
Keeping the structure, changing everything else
Structural conditioning feeds the model a second input alongside your prompt: a depth map, an edge trace, a pose skeleton, a rough scribble. The model must respect that structure while the prompt decides everything else.
This is how you get the exact same room at night in the rain, or a pencil sketch turned into a finished illustration that keeps your composition. It is the single biggest jump in control that most people have never tried.
Instruction editing, and its quiet problem
A newer class of model edits from plain instructions with no mask at all. "Make it evening." "Remove the man on the left." "Put her in a green jacket." These are genuinely good and have become the default for casual work.
Their weakness is worth knowing. Most of them re-render the whole image, not just the part you named. One edit is invisible. After five sequential edits, the face has drifted, the sign on the wall has changed its wording, and fine texture has softened toward the model's house style. It is a photocopy of a photocopy.
So keep the original. Do the biggest change first. Where you can, run independent edits from the same original and composite them, rather than stacking edits on edits.
Upscaling is not detail recovery
An upscaler invents plausible detail. It does not recover detail that was never captured. A good one turns a soft image sharp and convincing, and it will also confidently invent the wrong lettering, the wrong eyelashes, and a fabric pattern that was not in the shot. Upscale last, look closely at anything that carries meaning, and expect one more round of inpainting afterwards.
The mindset shift
Treat generation as photography, not as ordering. You take a lot of frames, choose one, then do the darkroom work. The people producing consistently good AI images are not writing better prompts than you. They are doing four rounds of masking you never see.
Before you move on