Two ways in, and they are not equal
Text-to-video: you type a description and the model invents the framing, the faces, the wardrobe, the colour, the light and the lens, then animates its own invention. You are gambling on a dozen decisions simultaneously, at video prices.
Image-to-video: you supply the first frame. The model's only job is motion.
Almost everyone doing this professionally works the second way, and the reason is economic before it is aesthetic. A still costs a cent or two and takes seconds. An eight-second clip costs fifty to a hundred times that and takes a minute or two in a queue. Image-to-video moves every hard decision out of the expensive loop and into the cheap one.
There is a second reason that matters more. A still is not only cheaper to regenerate — it can be repaired by hand. You can photograph it. You can composite the client's real product into it in Photopea, which runs in a browser tab, opens PSD files and needs no install. You can fix a bad hand in GIMP or Krita. You cannot do any of that inside a video model. Whatever you fix in the still is fixed for the whole clip.
For making the still, see ai-images. For repairing it, image-editing.
The first frame is a contract
Everything in that frame persists and then degrades. Six fingers in the still means six fingers for eight seconds, getting stranger. A wobbly sign in the still becomes a melting sign. A face you are not quite happy with in the still is a face you will hate by second five.
The frame also fixes the style. The model will not restyle mid-clip — it animates what it is given. So the still is where you decide whether this is a photograph, a gouache painting or a 3D render, and that decision is then locked in a way no prompt word can override.
The practical rule: do not animate a still you would not be happy to publish as a still.
Match the aspect ratio and the resolution
A quiet, common waste of money. You compose a still you like at 3:2. The video model works at 16:9 and either crops it, pads it or stretches it, and the framing you chose is gone.
Generate the still at the ratio the video model outputs — 16:9 for landscape, 9:16 for vertical short-form, 1:1 where a platform wants it — and at or a little above the model's native resolution. Most generate at 720p or 1080p. Feeding in a 4000-pixel still does not buy you detail; it gets downscaled on the way in.
Start and end frames
Several tools let you supply both the first and last frame — Luma's keyframes, Kling's start-and-end frames, and similar controls elsewhere. This is the most underused control in the whole category. It lets you decide where the shot lands, which is what a real camera operator does.
Use it for a reveal, for a match cut into the next shot, or for arriving exactly on a product hero frame.
The failure mode is worth knowing: if the two frames are too different, the model does not invent a camera move between them, it invents a morph. Keep them the same scene, same light, same lens, with a plausible path between. A push in works. A cut from a kitchen to a beach does not.
What text-to-video is still good for
It is not useless. When nothing is decided yet, text-to-video is a fast way to find a look, generate options and see what the model is good at before you commit. Concept phase, not production phase.
There is also a real trade-off on sound. Some models generate synchronised dialogue and ambience along with the picture, and on several of them that native audio path is tied to text prompting rather than to an image you supply. If you need the model's own audio, you may have to give up the control of a fixed first frame. Decide which matters more for that shot.
The free path
- Stills: several image tools give a daily free allowance, and open models run in a browser through a Hugging Face Space with no install. See hugging-face.
- Fixing stills: Photopea in a browser handles layers, masks and PSDs on a laptop that cannot run Photoshop. GIMP and Krita if you can install. On a phone, Snapseed for tone and a browser tab for the rest.
- Video: most hosted tools give a small allowance that refreshes. Spend it on image-to-video, at the lowest resolution, on real shots from your list. You learn far more per free credit that way than by rolling text-to-video prompts.
Today: take one shot you want. Make the still. Fix one thing in it by hand — the label, the hand, the sign. Then animate it, and compare that clip to one generated from text alone.
Before you move on