Key Takeaways
- Fractile spent its first two years building on-chip SRAM chips like Groq and Cerebras before abandoning the architecture entirely.
- SRAM breaks down under modern inference demands because both parameter counts and context lengths are expanding simultaneously.
- Long-running reasoning agents require multi-trillion parameter models running at thousands of tokens per second, which SRAM cannot support economically.
- Fractile partnered directly with memory foundries to design a DRAM architecture that delivers 25x the bandwidth of standard High Bandwidth Memory (HBM) without losing capacity.
The Two-Year Pivot Away From SRAM
For two years, Walter Goodwin built Fractile around the same technical bet as Groq and Cerebras: on-chip Static RAM (SRAM). The physics behind that choice seemed obvious at the time. SRAM sits directly next to compute cores, eliminating off-chip data transfers and serving tokens at blistering speeds.
“For the first kind of two years or so of the company's life,” Goodwin said, “we were working on, a little like Groq or Cerebras, working on an SRAM-based chip.”
Then the workload shifted. AI models did not just get wider across their weights; their memory footprints exploded along the context dimension. Reasoning models and autonomous agents started ingesting hundreds of thousands of tokens of context history, codebases, and tool calls.
SRAM provides extreme speed, but it has terrible physical density. Fitting a frontier model across SRAM requires chaining hundreds of separate silicon dies together. When you expand context windows, the key-value cache expands with it, demanding vast memory capacity that SRAM cannot economically supply.
“There are two things that grow with AI today,” Goodwin noted. “One is obviously the parameters of the model. But the other, and this was the one that kind of really got us nervous about that architectural approach, was this kind of growing context length that was becoming more and more part of the story for how we saw these models rolling out.”
Solving the Memory Bandwidth Bottleneck
When Goodwin realized SRAM would price Fractile out of frontier inference workloads, he did not fall back to standard off-the-shelf GPU architectures with standard HBM. Standard HBM chips still choke on memory bandwidth during the autoregressive generation phase, where every generated token requires streaming every weight from memory.
Instead, Fractile partnered directly with memory foundries to engineer a custom DRAM architecture. The goal was simple: match the density and capacity of DRAM while delivering 25x the memory bandwidth of standard HBM.
Goodwin sees raw bandwidth as the gating factor for autonomous agents. If an agent needs to execute hundreds of internal reasoning steps, code executions, and verifications before returning an answer, slow generation speeds make the system unusable. To make multi-step reasoning viable, hardware must run trillion-parameter models at thousands of tokens per second without running out of memory.
As Goodwin put it, the entire challenge reduces to a physical requirement: “Like many technical challenges, it boils down to a slightly mundane technical observation, which is we need chips that have this kind of particular ineffable property, which is incredibly high bandwidth to memory.”
What to Do With This
Audit your model inference stack this week by calculating the ratio between your active parameter footprint and your peak key-value cache size across your longest user sessions. If context memory exceeds 30% of total memory allocation, stop optimizing compute kernels and benchmark high-bandwidth memory alternatives, because your latency floor is dictated entirely by memory bandwidth.