Key Takeaways
- Cloud language models hide hardware inefficiencies by batching hundreds of concurrent user requests; physical robots cannot aggregate requests and must run batch-size-one inference.
- Autoregressive transformer models are memory-bound, meaning embedded edge chips starve compute units while waiting for memory bandwidth.
- Physical safety and real-time response times require processing on spot rather than routing sensor data to centralized data centers.
- The benchmark for physical AI hardware changes from raw throughput to intelligence per watt.
The Batch-Size-One Reality in Physical AI
Software engineers building on web APIs assume compute can always be batched. If ten thousand users send prompts to an LLM provider, the provider groups those prompts into large compute batches. This keeps the GPU compute cores saturated and amortizes model weights across requests.
Physical machines do not have that luxury. When a robotic arm reaches for a moving part on an assembly line or an autonomous system avoids an obstacle, it processes a single stream of sensor data in isolation.
Ming-Yu Liu, Vice President of Cosmos Lab at NVIDIA, points out that physical systems operate under hard constraints: “A lot of tasks is a batch size one inference problem. So API, you can aggregate users' API call from different users, and then process them in whole.” On a robot, that aggregation disappears.
Memory Bandwidth Breaks Autoregressive Models at the Edge
The architectural problem comes down to memory bandwidth. Large language models built on autoregressive transformer architectures generate tokens sequentially. Each generated step requires moving the entire model weights through memory just to process a single token.
In the cloud, batching masks this memory bottleneck by reusing loaded weights across dozens of queries simultaneously. When you move that exact model onto an embedded board inside a robot, the silicon stalls.
Liu explains the hardware disconnect directly: “LM is based on the autoregressive transformer. So it's memory bound. And to use the GPU power more efficiently, you can aggregate the API call from different user to fill in the computes. The same architecture if you put in this embedded device, you don't have API call to aggregate.”
Without thousands of calls to batch together, the embedded GPU spends most of its clock cycles waiting for data to travel across the memory bus. You burn battery, generate heat, and get terrible compute utilization.
Real-Time Latency Demands Intelligence Per Watt
Some teams try to bypass edge compute limits by offloading robotic reasoning to the cloud. That shortcut fails under physical operating realities. A robot operating around humans cannot pause for network jitter or dropped packets.
“The robot need to react very fast,” Liu notes. “So the processing have to be done real time on spot. There's a real time constraint and safety constraint. So, because of the setup difference, it gonna promote different architectures.”
Because compute must sit directly on the robot, power draws dictate the entire system design. A heavy compute payload drains batteries, demands bulky cooling systems, and limits payload capacity. Raw parameter counts become useless if the model draws two hundred watts on a battery-constrained chassis.
Liu states the core metric clearly: “In the end, the primary evaluation would be intelligence per watt.” Physical AI requires architectures designed specifically for single-stream, low-power execution rather than shrunken copies of cloud chat models.
What to Do With This
Audit your model deployment pipeline this week. Benchmark your model latency, memory bandwidth saturation, and power consumption strictly at batch size one on your target embedded hardware. If your architecture relies on batching or cloud round-trips to achieve acceptable latency, replace the model architecture before finalizing your hardware design.