Key Takeaways
- Closed APIs block internal inspection: Ming-Yu Liu points out that hosted endpoints prevent engineers from modifying model architectures or running local, zero-latency inference.
- Open models replace static code libraries: Liu frames model weights as the modern equivalent of software packages that developers pull down, adapt, and run locally.
- Frontier labs cannot predict edge use cases: Open weights allow developers to build specialized physical AI applications that base foundation model creators never planned for.
- Model workloads dictate silicon design: Releasing open models like Cosmos gives NVIDIA direct data on actual developer compute bottlenecks, directly steering future GPU architectures.
Closed APIs Cap Engineering Freedom
When a company builds on proprietary API endpoints, it trades control for convenience. You get a working result, but you inherit an intellectual wall. You cannot inspect internal attention weights, you cannot run post-training on edge devices, and you cannot run workloads without an active network connection.
Ming-Yu Liu, Vice President of Cosmos Lab at NVIDIA, argues that real engineering demands deeper access. "Innovation freedom," Liu explained. “You can use an API to solve the problem, but with an open model, you can do more than what is available in an API.” If you cannot open the engine, you cannot tune it for physical simulation or robotic control.
“APIs are great. They solve problems, but they don't give you the insight,” Liu noted. “It doesn't allow you to tear apart and mix your idea inside.” For teams building physical AI, real-time control loops and unique sensor arrays cannot wait on a remote server call.
Models Are the New Software Libraries
Software development historically relied on open-source libraries. If a standard linear algebra package lacked a specific matrix operation, engineers patched the code directly. Liu argues that machine learning models now occupy that exact structural role in engineering pipelines.
“In the modern world, you have computers and you have libraries still, but you now have models,” Liu said. “Model as a new kind of library, and people can use them to build great, amazing applications.” When models act as libraries, developers treat weights as raw code rather than finished products.
This matters because frontier AI labs cannot predict what specialized builders need. “There will be applications that were not foreseen by the foundation model frontier companies,” Liu stated. An industrial automation team training a robot arm to handle reflective sheet metal requires custom architectures and domain-specific post-training. A closed endpoint offers no path to modify those internal representations.
The Hardware Feedback Loop
NVIDIA does not release open models out of charity. It does so to observe how developers break, modify, and run workloads under production constraints. When thousands of engineers fine-tune open weights on real-world datasets, their compute patterns show NVIDIA where current hardware falls short.
“Through building the model, we better understand what kind of computer we should build, what kind of GPU architecture we should have,” Liu explained. Releasing models like Cosmos creates a direct feedback mechanism: developer experimentation reveals memory bandwidth limits, latency bottlenecks, and quantization trade-offs, which NVIDIA then uses to design its next generation of chips.
What to Do With This
Audit your core product pipeline this week. Identify any production feature that depends on a closed API where you cannot control latency, uptime, or custom weight tuning. Replace that single endpoint with an open-source model hosted on your own infrastructure, and test whether local post-training improves your task accuracy.