Key Takeaways
- The sheer scale of AI inference is so vast that it naturally drives the creation of specialized Application-Specific Integrated Circuits (ASICs), distinct from general-purpose GPUs.
- New chip companies like Etched aren't aiming to unseat Nvidia's dominance in AI training; instead, they are evolving the landscape set by previous inference specialists like Cerebras and Groq.
- Etched's designs are optimized for current, dominant AI architectures, specifically targeting post-transformer and RGBT models, which differ from older, general-purpose approaches.
- While AI model architectures constantly shift, the widespread and sticky use of established models like GPT-3.5 provides a stable, multi-year market for specialized inference chips.
The Inference Tidal Wave: Why Nvidia Can't Handle It All
Founders often default to Nvidia for AI workloads, and for good reason—their GPUs are the gold standard for training. But Swyx, founder of the AI Engineering Conference (AIE), argues that focusing solely on Nvidia misses a massive, emerging opportunity: AI inference. “All of inference is just so goddamn big,” Swyx explains, “that of course you're going to have ASICs for inference.” The logic is simple: when a computational task reaches mind-boggling scale, generic hardware becomes inefficient. Think about Bitcoin mining or video encoding; eventually, dedicated chips emerge that do one thing extraordinarily well. AI inference, from powering ChatGPT queries to running image generators, is reaching that point.
This isn't about disrupting Nvidia's core training business. It's about recognizing that the compute needs for running AI models are fundamentally different and larger than building them. Training is intense, bursty, and demands flexibility. Inference is high-volume, repetitive, and screams for efficiency. This distinction opens the door for companies like Etched to carve out a critical niche, not by competing head-on with Nvidia, but by optimizing for the unique demands of serving billions of AI requests.
Etched: Not an Nvidia Killer, but a Cerebras Upgrade
So, if not Nvidia, who are these new chipmakers after? Swyx offers a clear framing: think of Etched as a "next generation Cerebras, next generation Groq." Cerebras, for example, started over a decade ago, pioneering massive, specialized chips. Their early designs were impressive, but the AI landscape has evolved. What Etched, Maddx, and other new players are doing, Swyx notes, is “much more dedicated to like basically everything that we know… post transformers, post RGBT and optimizing for that workload.” This means their designs bake in assumptions about the fundamental structure of modern AI models, allowing for far greater efficiency than older, more general-purpose ASICs.
It's a subtle but critical distinction. These companies aren't trying to invent a new universal computing paradigm. They're taking the widely accepted, high-impact architectures of today—like the transformer model that powers most LLMs—and building hardware perfectly suited for them. This focus allows them to squeeze out unprecedented performance per watt and dollar, making inference cheaper and faster for the applications that truly matter right now.
The GPT-3.5 Bet: Stability in a Shifting Landscape
One might worry that building ASICs for specific model architectures is a risky game. What if the next big breakthrough renders today's chips obsolete? Swyx counters this with a pragmatic observation about how AI models are actually used in production. “The chat GPT like GPT 3.5ish architecture has mostly stayed the same this entire time,” he points out. “It's a pretty good bet, man.” While research continues to push boundaries, the models widely deployed today—the ones doing the heavy lifting for millions of users—tend to stick around because they simply work.
Founders and engineers don't rip out working systems lightly. “Once the thing works, it works. Like don't don't touch it,” Swyx says. This inertia creates a long tail of demand for specific model architectures. So, even as GPT-5 or other frontier models emerge, the massive inference workload for current, stable models like GPT-3.5, GPT-4, or specific vision transformers will persist for years. This stability makes building specialized hardware a viable, even lucrative, long-term play for companies like Etched.
What to Do With This
If you're building an AI application, stop defaulting to general-purpose GPUs for every inference task. This week, investigate specialized inference providers or hardware for your primary models. If your application relies heavily on stable, widely used architectures (like transformer-based LLMs or common vision models), you might find significant cost savings and performance gains by using ASICs tailored for those workloads, rather than general-purpose compute that adds unnecessary overhead.