12 quotes from 1 episode on Practical AI, each with a timestamped link to the source.
12 quotes1 episode
The short version
Ming-Yu Liu argues that physical AI represents the next major market for hardware because robots interacting with reality run local, instantaneous processing. Creating world models and open software frameworks directly shows NVIDIA how to design future GPU architectures to meet these on-device demands.
Most interesting insights
Model weights serve as the modern equivalent of software packages that engineers pull down and adapt.
“In the modern world, you have computers and you have libraries still, but you now have models…”
Ming-Yu Liu, Practical AI · October 2026 · Listen ↗
Cloud applications efficiently process hundreds of aggregated user requests at once. Physical machines operate alone and process commands one at a time.
“A lot of tasks is a batch size one inference problem. So API, you can aggregate users' API call from different users, and then process them in whole.”
Ming-Yu Liu, Practical AI · October 2026 · Listen ↗
Constructing foundation models provides direct feedback on compute limitations. Understanding these constraints shows NVIDIA exactly what kind of GPU architectures developers demand.
“Through building the model, we better understand what kind of computer we should build, what kind of GPU architecture we should have…”
Ming-Yu Liu, Practical AI · October 2026 · Listen ↗
A complete open release provides developers with specific data sets and software frameworks. Engineers use these tools to train base models on proprietary information.
“Open model, it's not just have the model open weight allow you to use. We also provide the training framework, so that you can take this open model and your own data, and post-train to something more tailored for your use case.”
Ming-Yu Liu, Practical AI · October 2026 · Listen ↗
Simulated world models generate pixel-level environmental changes. These environments provide autonomous systems with exact control signals to complete physical tasks safely.
“With a world model, instead of having your policy deploy the real car, drive in the real world, you can have your policy interact with the world model…”
Ming-Yu Liu, Practical AI · October 2026 · Listen ↗
“When you model the world, it generates the pixel space evolution, and the pixel space evolution has a strong correlation to the control signal you might need to use to complete certain manipulation tasks.”
Ming-Yu Liu, Practical AI · October 2026 · Listen ↗
“If we can make physical AI come, you know, the the dream of physical AI come true, there will be a lot of opportunity for NVIDIA, so we are very aligned…”
Ming-Yu Liu, Practical AI · October 2026 · Listen ↗
Closed APIs block internal inspection: Ming-Yu Liu points out that hosted endpoints prevent engineers from modifying model architectures or running local, zero-latency inference.
Open models replace static code libraries: Liu frames model weights as the modern equivalent of software packages that developers pull down, adapt, and run locally.
Physical AI refers strictly to models deployed on physical devices that perturb the physical state of the environment to produce tangible, material results.
Ming-Yu Liu identifies four core verticals driving physical AI: autonomous vehicles, factory automation, agriculture, and humanoid robotics.
Cloud language models hide hardware inefficiencies by batching hundreds of concurrent user requests; physical robots cannot aggregate requests and must run batch-size-one inference.
Autoregressive transformer models are memory-bound, meaning embedded edge chips starve compute units while waiting for memory bandwidth.
Ming-Yu Liu, Vice President of Cosmos Lab at NVIDIA, defines an open model by three components: open weights hosted on Hugging Face, training frameworks on GitHub, and shared datasets.
The open release approach spans three NVIDIA lines: Cosmos for physical simulation and world models, Nemotron for language, and Project GR00T for robotics.
How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode, and we use it only when a separate check of the captions finds that person on the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.