Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Stop describing it and show it the picture

Image Generation, In Practice · lesson 3 of 10 · 8 min

Adjectives are a bad way to specify a layout

You want a shot of your uncle's tailoring shop: the counter on the left, the fabric rolls stacked to the ceiling on the right, the doorway blown out with daylight. You can write four hundred words about it and the model will still put the counter wherever it likes, because you are describing a picture to something that matches captions rather than reads floor plans.

You have a phone. Take the photograph. Then hand the photograph to the model.

This is the hinge of the whole course. Almost every technique from here on is a way of giving the model something other than words.

Denoise strength is the only dial that matters

Image-to-image works like this: instead of starting from pure noise, the model starts from *your* image with a controlled amount of noise mixed in, then denoises from there. How much noise it adds is the strength — sometimes labelled denoise, denoising strength, or image weight.

That one number decides how much of your picture survives.

  • 0.15-0.25 — your image back, slightly repainted. Useful for a texture or style tweak, useless for changing content.
  • 0.4-0.6 — the working range. Same composition, same masses of light and dark, genuinely different rendering. This is where most useful img2img lives.
  • 0.75 and up — the model has effectively been handed static. You get something loosely inspired by your image and you have wasted the reference.

People try 0.8, see nothing of their photo, conclude img2img does not work, and go back to writing adjectives. The range is narrow and you find it by bisecting: run 0.3, 0.5 and 0.7 on the same seed, look at all three, then split the interval that looked closest.

Three different things called "reference"

Tools confuse these and it causes real frustration.

  • Structure reference — keep the layout, change everything else. This is img2img at mid strength, and it is what you want for the tailoring shop.
  • Style reference — take the palette, the brushwork, the grade, and apply it to a different subject. Midjourney exposes this as --sref; hosted tools usually call it a style image.
  • Character or subject reference — carry a specific face or object into a new scene. Midjourney's --cref, and the identity features covered in lesson six. This is the hardest of the three and the one most likely to disappoint.

If a tool gives you one box marked "reference image", find out which of the three it is doing before you blame your input.

What ruins a reference

Resolution and aspect ratio. A 400-pixel WhatsApp forward carries compression blocks, and at mid strength the model will faithfully reproduce blockiness as texture. Shoot at full resolution. Match the reference's aspect ratio to your output, or the model crops or stretches in ways you did not ask for.

Busy backgrounds. The model has no idea which parts of your reference you care about. If the fabric rolls matter and the pile of boxes does not, crop the boxes out before you upload. A cleaner reference is a stronger instruction.

Text in the reference. It will come back as almost-words. Cover it or crop it, and add real type later.

Where this leaves you, and where it does not

Img2img gives you a composition you chose rather than one the model guessed. It does not give you precise control: at 0.5 strength the doorway is still in the doorway's place, but the number of fabric rolls will change and the counter's proportions will drift. If you need the geometry to hold exactly — a pose matched frame to frame, a room's perspective preserved — that is conditioning, and it is the next lesson.

Free path, all of it: Krita with the AI Diffusion plugin gives you a canvas, a strength slider and local generation, free. ComfyUI and InvokeAI the same. On a phone, most hosted tools accept an uploaded image in their app, and the Gemini app in particular is built around editing a picture you give it.

The consent line, briefly

A style reference taken from one living artist's work, used to produce commercial output in their manner, is not a technical setting. It is a decision about somebody's livelihood, and courts in several countries are currently arguing about it. Referencing a movement, an era or a printing process is a different act from referencing a named person still selling work. Lesson ten returns to this.

Today: photograph something in your own room, run it through any img2img tool at 0.3, 0.5 and 0.7 with the same prompt and seed, and put the three results side by side. You now know that dial for that tool, which is knowledge no article can give you.

Before you move on

Someone feeds a phone photo of their shop into an image-to-image tool at 0.85 strength and gets a picture with nothing of the shop in it. They drop to 0.15 and get their own photo back with faint grain. What is going on?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly