Key Takeaways
- Over the past 20 years, raw compute FLOPs scaled 1,000,000x, while memory bandwidth scaled only 40x.
- Frontier AI labs want to push Mixture-of-Experts (MoE) architectures from standard 1-in-16 routing to extreme sparsity like 1-in-128 or 1-in-256 routing.
- Current High-Bandwidth Memory (HBM) GPUs cannot serve hyper-sparse architectures efficiently because memory bandwidth starves the compute cores during token generation.
- Fractile targets 25x more memory bandwidth per chip than standard HBM chips to slash total FLOP requirements at matched intelligence levels.
The Million-Fold Compute Imbalance
AI hardware development has spent two decades running in one direction: packing raw arithmetic compute onto silicon. That single-minded push created a massive architectural debt.
“We've scaled flops like a millionfold in the last 20 years,” Walter Goodwin explains. “Memory bandwidth has gone up about 40x in the same time frame.”
Compute speed outpaced memory access speed by a factor of 25,000. When training dense transformer models, this imbalance was manageable because dense matrix multiplications reuse weights many times per memory load. Inference generation is different. Generating tokens one by one requires reading model weights out of memory for every single step. Compute cores sit idle while waiting on memory buses.
Why HBM Chokes on Sparse MoE Models
Mixture-of-Experts architectures offer a clear mathematical path to cheaper intelligence. Instead of activating every parameter for every token, the model routes each token to a tiny fraction of its total expert networks.
Frontier labs know the math favors extreme sparsity. As Goodwin notes: “These mixture of expert models, it's pretty well known that actually ideally we would make them sparser and sparser and sparser. So for like ISO intelligence, you will save a ton of flops if you go from being like 1 in 16 sparse on your to 1 in 128, 1 in 256.”
Under current hardware constraints, that theory falls apart in production. “One of the challenges, one of the headwinds to doing that is actually that it becomes incredibly prohibitive on today's HBM based GPUs, XPUs to serve those models efficiently. You end up often bandwidth bottlenecked,” Goodwin says.
When a model routes to 1 out of 256 experts, it still needs fast access to vast pools of distributed memory to retrieve the active weights. HBM does not provide enough bandwidth per chip to pull those weights without stalling the pipeline. Chip designers are forced to run denser, less efficient models simply because today's hardware cannot feed sparse ones.
Fractile is building around this exact gap. By delivering “25 times more bandwidth per chip than an HBM-based chip,” hardware can finally match sparse routing algorithms. When memory bandwidth scales, developers can drop FLOP counts sharply while maintaining model quality. Goodwin points out that this relationship creates a multiplier on global throughput: higher bandwidth directly reduces the compute needed to reach a target intelligence threshold.
What to Do With This
Audit your model inference economics this week. Profile your production token generation to measure your arithmetic intensity and determine whether your hardware is memory-bandwidth bound or compute bound. If you are training or fine-tuning MoE models, calculate the serving cost tradeoffs of moving to sparser expert routing (such as 1-in-64 or 1-in-128) across bandwidth-optimized hardware clusters.