Key Takeaways

  • Ming-Yu Liu leads NVIDIA Cosmos Lab to build foundation models that simulate physical physics rather than parse web text.
  • Cosmos 3 processes text, audio, video, and action control data as both inputs and outputs across every permutation.
  • World models replace risky real-world robot trial-and-error by generating future pixel dynamics directly from proposed actions.
  • Predicting pixel space evolution gives robotic systems the exact control signals needed to complete real-world manipulation tasks.
  • Simulation and understanding share a single unified representation inside Cosmos 3 because physical reality obeys one set of rules.

Text LLMs Hallucinate Words; Physical AI Breaks Hardware

Training an autonomous vehicle or robotic arm in the real world is expensive, dangerous, and slow. If an autonomous driving policy makes a bad decision on a public road, the result is a physical crash.

Ming-Yu Liu points out that large language models cannot solve this problem because text tokens ignore real-time sensory physics. A world model operates from sensory observations, capturing how cameras see the environment and how physical laws dictate movement over time.

“With a world model, instead of having your policy deploy the real car, drive in the real world, you can have your policy interact with the world model,” Liu explains. “When you steer left, what are you going to see? Steer right, what are you going to see?”

By predicting sensory outcomes before hardware moves, developers test control policies against neural simulations rather than risking bent metal.

Merging Physical Understanding and Simulation in Cosmos 3

Previous robotics stacks treated world simulation and environmental perception as separate software pipelines. Liu argues this separation created friction and errors.

In Cosmos 3, NVIDIA merged these layers into a single foundation model architecture. “In COSMOS 3 our latest COSMOS model, we actually fuse them together, because we believe that the representation can be shared, because we live in the same world,” says Liu. "It's the same stuff."

This shared representation supports text, audio, video, and action as symmetric inputs and outputs. A developer can feed a robotic arm a video stream and a natural language instruction, then receive motor control actions. Alternatively, the developer can feed motor actions and receive simulated future video frames.

The connection between visual frames and motor actions turns out to be tighter than previously assumed. As Liu notes, “When you model the world, it generates the pixel space evolution, and the pixel space evolution has a strong correlation to the control signal you might need to use to complete certain manipulation tasks.”

The Shift from Static Teleoperation to Omni-Modal Training

Most robotics startups still depend on expensive human teleoperation to gather physical trajectories. That approach caps your dataset at the number of physical hours humans spend wearing VR headsets or guiding robot wrists.

World models flip this bottleneck. Once a neural model internalizes the physics of how objects roll, drop, collide, and bend, it generates synthetic sensory streams for training policies at software speed. You move from collecting thousands of manual demonstrations to generating millions of simulated physical outcomes across varying lighting, friction, and object textures.

What to Do With This

Audit your robotics or autonomous system data pipeline this week. Identify where your team spends budget collecting manual hardware demonstrations. Swap one real-world edge-case test loop for a world model rollout where actions predict future visual frames before running on physical hardware.