Key Takeaways
- Text models cannot learn low-level physical dynamics because human language fails to describe subconscious visual interactions.
- Predicting raw video frames forces neural networks to simulate mechanics, fluid dynamics, and optics without hand-coded rules.
- Runway tracks physical comprehension across Gen-1, Gen-2, and Gen-3 using an empirical benchmark called PhysicsIQ.
- Scaling compute alone predictably boosts PhysicsIQ scores, showing that video scaling laws operate like language model scaling laws.
The Disagreement
AI researchers are split on how machines will understand the physical world. Skeptics like Yann LeCun argue that generative video diffusion models merely produce surface-level pixel illusions. In this view, predicting the next frame in a video cannot teach a model causal reasoning or real physical mechanics. Skeptics claim that building true world models requires joint-embedding architectures rather than generative pixel matching.
Germanidis takes the opposite side. He argues that predicting raw video pixels at scale forces a network to build an internal simulator of physical reality. As Germanidis puts it, “in order to predict video well you need to simulate the world in an increasing and increasing capacity and if scaling laws apply on video just like they apply on language models then as we scale the computer that we put in those models, then they're going to be able to simulate physics.”
Language models cannot fill this gap because human writing rarely describes basic mechanics. “There is just so much complexity and detail in the world that in order to that it's it's hard to learn directly from just human descriptions of the world,” Germanidis explains. “We're constantly underestimating all the complexity that goes into very like things that we do subconsciously as humans and we don't even necessarily always have the words to describe them.”
Who's Right (and When They're Wrong)
Germanidis has empirical evidence on his side. Runway measures how well models understand real-world dynamics using an internal benchmark called PhysicsIQ. As Runway scaled from Gen-1 through Gen-2 to Gen-3, the models did not just output higher-resolution frames. Their physical accuracy jumped in direct proportion to compute. As Germanidis notes, “we see that the score and physics IQ predictably improves. There's other you know tricks and techniques that you can make to improve the score even further but even scale alone helps in in the in the model learning better physics.”
The skeptics are wrong about the limits of scaling pixels, but they are right about efficiency. Predicting every single photon in high-resolution video requires massive compute clusters that few startups can afford. Scaling alone makes the physics score go up, but raw compute is an expensive teacher when modeling rare edge cases or precise contact mechanics for robotics.
What to Do With This
Stop trying to teach physical domain knowledge to models through text prompts and manual logic trees. If you build AI products that touch physical environments, evaluate your models against visual simulation benchmarks like PhysicsIQ rather than text-based logic tests. Run an evaluation this week where you compare your model's outputs against video ground truth across five distinct physical interactions: gravity, fluid flow, object collisions, friction, and light refraction.