Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Video, and what it still cannot do

Making Things With AI · lesson 5 of 8 · 8 min

The extra dimension costs more than you think

An image model has to make one frame plausible. A video model has to make 150 frames plausible *and* mutually consistent. It cannot make them one at a time, because frame 80 has to know what frame 12 did. So it compresses time as well as space — a video autoencoder might squash eight times in each spatial direction and four times in time — and a transformer attends across that whole compressed block at once.

Attention cost grows steeply with the size of the block. That is the real reason clips are short. Five to ten seconds is where quality, price and attention length currently meet. It is a physics-of-the-method limit, not an arbitrary product decision.

What genuinely works now

Single shots up to about ten seconds, with a clear subject and a simple camera move, look good. Image-to-video — where you generate or photograph a still first, then animate it — is far more controllable than text-to-video, and it is what most working people use. Several models now produce synchronised dialogue and ambience along with the picture, which was science fiction two years earlier.

What still fails

Physics at the point of contact. Anything touching anything else. Hands gripping, feet meeting ground, liquid pouring, cloth folding, a knife going through a tomato. The model learned what these look like, not what they are, and contact is where the difference surfaces.

Object permanence. When something leaves frame and comes back, it comes back subtly wrong. Crowds are the worst case; background faces churn continuously.

Identity across shots. Your character in shot three is a cousin of your character in shot one. Locked keyframes and reference features help. They do not solve it.

Timed action. "She looks up on the fourth beat" is not currently something you can ask for. You get an approximation and cut around it.

Text. Signage in video is roughly where image models were in 2022.

Length. You do not get a film. You get shots.

The number that matters is cost per usable second

List prices in 2026 sit roughly between ten cents and seventy-five cents per generated second, depending on model, resolution and tier. That number misleads, because what you actually spend is cost per second you can use. At a one-in-eight hit rate — normal once a shot has a specific requirement — an eight-second clip listed at two dollars really cost sixteen, plus an hour of your attention.

Budget for the takes you throw away. Teams that budget for the keeper run out of money at shot four of thirty.

A workflow that survives a deadline

  1. 1Write the shot list first, as shots, with durations. Treat it like a real shoot, because it is one.
  2. 2Make a still for every shot and get the still right. Image tools give you far more control than video tools do.
  3. 3Animate one shot at a time, several takes each.
  4. 4Assemble in an ordinary editor. Cuts hide an enormous number of sins. A long continuous move hides none.
  5. 5Do sound separately unless the model's native audio happens to land.

Where the honest line sits

For mood pieces, abstract backgrounds, product motion and B-roll nobody studies frame by frame, this is production-ready and cheap. For anything with a named character doing specific things in a specific order, you are fighting the tool, and a small live shoot or 2D animation is often faster, cheaper and better.

Say that in week one rather than week four. Being the person who knows where the tool stops is more valuable than being the person who can drive it.

Before you move on

A team plans a three-minute explainer as one continuous camera move through a factory, generated as a single piece. What is the strongest reason to change the plan?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly