Key Takeaways
- GPUs represent the single most expensive physical asset inside an AI data center; keeping compute engines idle destroys unit economics.
- Token generation slows down during inference when systems recompute attention matrices instead of reading previously computed state.
- Key-Value (KV) caches rapidly expand beyond on-chip High Bandwidth Memory (HBM), requiring tiering across system DRAM and NVMe drives.
- The provider that delivers the lowest cost per token will capture the largest market share as inference demand overtakes training runs.
Idle Silicon Burns Margin
When Crusoe closed a $3.9 billion Series F, CEO Chase Lochmiller focused on a simple reality of data center math: expensive chips cannot sit idle. In model training, compute runs flat out for months on static datasets. Inference is different. User traffic spikes unpredictably, prompt lengths vary wildly, and GPUs wait on memory transfers between token generations.
Winning in model serving does not come down to raw theoretical compute on a spec sheet. It comes down to token throughput per dollar. If your software pipeline stalls while waiting for weights or context to load into registers, your operating costs explode.
The KV Cache Memory Hierarchy
The primary bottleneck in transformer inference is attention memory. As a model processes a prompt, it generates keys and values for every token to track context across the sequence. Recomputing these values for each subsequent token wastes valuable processor time.
Lochmiller pointed to this exact mechanism: “When tokens are fed into a large neural network you can compute the output tokens by basically running this feed forward process and running all these matrix multiplications. That takes time. You may also already know the answer of that output token.”
Caching these calculations solves the compute problem but creates a storage capacity crisis. “The KV cache can get quite large and so it can expand well beyond the amount of memory you have on chip in the HBM,” Lochmiller explained.
High Bandwidth Memory (HBM) attached directly to the accelerator is fast but scarce. An Nvidia H100 accelerator offers 80GB of HBM3 memory. Long-context requests and concurrent user sessions quickly fill that space. Lochmiller stressed that managing this data flow across layers dictates serving efficiency: “Being able to manage that both across multiple different GPUs HBM layers as well as system memory DRAM as well as things like NVMe is actually a critical aspect to being able to serve inference and be able to keep those GPUs busy.”
Efficient inference engines do not just throw more accelerators at long context windows. They tier the KV cache. They hold active tokens in HBM, spill dormant prefix caches into host system DRAM over PCIe buses, and page cold context out to local NVMe drives. When you avoid redundant matrix multiplications without blowing your HBM budget, your token serving margin turns positive.
What to Do With This
Audit your model serving latency profile this week. Profile your average time-to-first-token against inter-token latency across your production clusters. If your GPUs show sub-50% compute utilization while memory sits pegged at 95%, implement prompt caching and offload your static system prompts to host DRAM rather than reserving raw HBM for duplicate context.