Testing a claim that it understands physics
The claim, and why it is attractive
Video models are increasingly described as world models or world simulators — systems that have learned how things behave, not merely how they look. The claim matters commercially, because a model that understands physical consequence is a different product from a model that generates pretty footage.
It is also, in a weak form, plausible. To predict the next frames of a falling object you must produce acceleration that looks right. Something about the dynamics has been captured.
The question is what "captured" means, and it is answerable by test rather than by argument.
Four tests that separate the claims
Occlusion and permanence. Show a distinctive object passing behind a larger one and re-emerging. A model with an internal object representation returns the same object. A model doing local plausibility returns an object of roughly that kind, or a different one, or none. This is the single most informative test and it is easy to run.
Conservation. Pour liquid from one vessel into another and watch the volumes. Cut an object and count the pieces. Break a plate and see whether the fragments could reassemble. Quantity is a global constraint, and the counting lesson tells you what to expect.
Causal order. Ask for the consequence of an event rather than the event: a glass that has already fallen, a candle after it was blown out, a footprint after the walker has passed. Models trained on captioned clips have seen far more of the event than of its aftermath.
Counterfactuals and unusual physics. A ball bouncing on a surface at an unusual angle, an object with the wrong weight, a sport played with the wrong equipment. If the model has learned dynamics it can extrapolate; if it has learned the appearance of familiar clips, it snaps back to the familiar version.
Run these on any model whose marketing makes the claim. It takes twenty minutes and it is more informative than any published benchmark score, because these are exactly the cases benchmarks under-represent.
What the tests currently show
Honestly reported: current strong models do well on short, simple, common dynamics — falling, splashing, cloth, smoke, walking. They fail with regularity on occlusion and permanence, on conservation of quantity, on causal aftermath, and on anything counterfactual.
That pattern is consistent with a model that has learned a very good statistical prior over how footage of the world looks over short spans, and inconsistent with one that maintains a representation of objects and their properties.
This is not a criticism of the models. A very good visual prior is enormously useful and is what most production work needs. It is a criticism of the word "understands", which does specific work in a sales conversation.
Why it matters beyond argument
Two practical reasons to care.
It predicts your failures. If the model has no object representation, you should expect exactly the problems this module has described, and you should plan shots that avoid them rather than hoping the next release fixes them. It has not fixed them for three releases.
It matters for downstream claims. World-model language is being used to argue that these systems can serve as simulators for robotics, driving and planning — domains where being wrong about occlusion has consequences. The claim may eventually be true. The tests above are the honest way to find out, and they are the ones a buyer should ask for.
A model that predicts what footage looks like is not the same as a model that predicts what happens. The two agree most of the time, which is exactly why the disagreements are worth finding on purpose.
There is a related test worth adding, because it separates memorised footage from learned dynamics more sharply than the physics cases. Ask for something that is physically ordinary and visually rare: a cricket ball rolling across a snooker table, a person walking backwards up stairs, a candle burning in a strong wind. Each is easy to imagine and unlikely to be well represented in any video corpus. A model with dynamics extrapolates; a model with a strong appearance prior returns the familiar version — the ball on a cricket pitch, the person walking forwards — or produces motion that reads as reversed footage rather than as the action requested.
The unsettled part, stated plainly: there is a genuine research disagreement about whether enough video training induces a physical model as a by-product, or whether an explicit structure is required. Serious people hold both positions. What is not in dispute is that current systems fail the tests above, and anyone claiming otherwise should be asked to run them.
The one thing to keep
Claims that a video model has learned physics should be tested with occlusion, conservation and causal-order cases, because local plausibility over a few seconds is achievable without any physical model at all.
Before you move on
Which test most directly distinguishes a model with an internal object representation from one producing locally plausible footage?
Pick the one you would defend. Nobody sees your answer.