Key Takeaways
- Global chip supply chains are running at 100% capacity with hardware shortages across every compute category.
- Cloud data centers cannot handle complete model execution alone, forcing software architecture to split workloads between local silicon and remote servers.
- Qualcomm acquired Modular to open-source its AI software stack across both data center chips and edge processors.
- Smart glasses represent the next major hardware shift because physical proximity to eyes, ears, and mouth enables real-time multimodal inputs.
The Hardware Bottleneck Is Everywhere
Every founder building software assumes compute will get cheaper and more abundant every quarter. Cristiano Amon sees the opposite side of that ledger. As head of Qualcomm, he watches chip fabrication plants operate flat out.
“There's so much more demand for computes than availability right now and across the board,” Amon explained. “As a semiconductor company I'll tell you the supply chain it's operating at 100% capacity. Everything, everything is short.”
This shortage kills the fantasy that companies can route every user interaction to massive cloud clusters. If your application sends every keystroke and audio packet to an H100 cluster in Virginia, your unit economics collapse under high traffic. The only architectural escape hatch is running smaller inference steps directly on phones, cars, and laptops.
Smart Glasses and the Local Compute Split
Software builders often debate whether AI models belong entirely on the device or entirely in the cloud. Amon considers that debate settled by looking at how phones already operate.
“Let's go have a conversation about what part of your app runs on the device or runs on the cloud,” Amon said. “Even though we build incredible processors, if you put on airplane mode you don't use your phone. AI is no different than that; it's going to be running on the device and on the cloud, it's going to be all transparent to you.”
The next step for this split is hardware worn on the face. Amon pointed to smart glasses as the coming consumer shift because they sit on prime human real estate.
“I actually a big believer that glasses is going to see an inflection point,” Amon said. “The glass is a prime real estate. Close to your eyes, to your mouth, to your ears. Your head turns, your camera see it, and then those things like see what I see, read what I read, hear what I hear are going to come up.”
Running vision models from head-mounted cameras over cellular networks drains batteries and adds latency. Orchestration must happen on the frame, handing off heavy reasoning to the cloud only when required.
Open Sourcing to Break Single-Vendor Lock
To prevent developers from getting trapped inside proprietary compute silos, Qualcomm bought Modular and made a sharp strategic decision: release the stack to the public.
“So we bought that company and we're making that open source,” Amon stated. “Because we actually believe that the industry will benefit from an open-source stack that scale from the data center across different hardware in the edge.”
When a single hardware supplier controls the compiler and the runtime, developers pay higher margins and lose architectural flexibility. By opening the software layer that bridges data center chips and mobile silicon, Qualcomm wants to ensure model weights compile cleanly across diverse processors without rewriting core inference logic.
What to Do With This
Audit your application's inference calls this week. Identify which tasks, such as text embedding, input validation, or basic classification, can run locally on client hardware via WebGPU or local runtimes instead of hitting paid cloud endpoints. Moving 20% of your prompt processing to user devices will drop your monthly server bills immediately.