11 quotes from 1 episode on No Priors, each with a timestamped link to the source.
11 quotes1 episode
The short version
Walter Goodwin argues memory bandwidth limits modern AI capabilities far more than raw compute power. Hardware developers scaled processing operations 1,000,000x over 20 years, yet current architectures still choke on highly sparse models.
Most interesting insights
Internal hardware programs at major technology firms often exist primarily to force volume discounts from external chip suppliers.
“There's a bit of a joke today that the sort of first party efforts their primary purpose is to reduce the price that people pay Nvidia.”
Many application-specific hardware designs currently on the market offer nearly identical architectures.
“…if you look across this entire space of AI ASICs, one of the things that is very striking is there is a lot of relatively identical chips out there.”
Standard memory bandwidth restricts advanced model architectures
Developers want to build highly sparse mixture-of-expert architectures routing 1 in 256 parameters. Standard graphics processing units cannot serve these designs efficiently because memory delivery starves the compute cores.
“These mixture of expert models, it's pretty well known that actually ideally we would make them sparser and sparser and sparser. So for like ISO intelligence, you will save a ton of flops if you go from being like 1 in 16 sparse on your to 1 in 128, 1 in 256.”
“One of the challenges, one of the headwinds to doing that is actually that it becomes incredibly prohibitive on today's HBM based GPUs, XPUs to serve those models efficiently. You end up often bandwidth bottlenecked…”
Proprietary silicon creates algorithmic vulnerability windows
A lab running entirely on custom hardware faces a 9-month delay if a competitor discovers a computational breakthrough. That delay occurs while waiting to deploy hardware capable of running the updated algorithms.
“Suppose I'm lab one and I've gone all in on some proprietary silicon and then lab 2 discovers some new computational breakthrough that delivers far better computational efficiencies for the same level of intelligence but it only works on the chip that they've decided to deploy. I could die in the 9 months before I get to deploy enough of that chip that I've also now gained that kind of 5x in computational efficiency.”
Hardware speed relies on cutting physical production delays
Foundries enforce a physical minimum of 3 to 5 months to manufacture and return a silicon design. Compressing the time between a technical observation and high-volume delivery captures market value.
“…from the moment you send the chip to them to getting it back is you know 3 to 5 months even in a kind of super hot lot scenario. And so these are the kind of innate latencies I guess in the industry.”
“…the shorter you can make that latency, that gap between an observation and realizing that bet in volume, that is where there's an, you know, an enormous amount of value to be captured.”
“Like many technical challenges, it boils down to a slightly mundane technical observation, which is we need chips that have this kind of particular ineffable property, which is incredibly high bandwidth to memory.”
Hyperscaler ASICs like Google TPU, Meta MTIA, Microsoft Maia, and OpenAI Jalapeno rarely offer differentiated computational performance over Nvidia or AMD.
Most custom chips exist primarily as commercial bargaining tools to force volume discounts on merchant silicon orders.
Physical fabrication enforces a strict floor of 3 to 5 months from tape-out to physical chip delivery, even when running expensive super hot lots at foundries.
Silicon economics require hardware to deliver a useful lifespan of at least 3 years to amortize development and production expenses.
Over the past 20 years, raw compute FLOPs scaled 1,000,000x, while memory bandwidth scaled only 40x.
Frontier AI labs want to push Mixture-of-Experts (MoE) architectures from standard 1-in-16 routing to extreme sparsity like 1-in-128 or 1-in-256 routing.
How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode, and we use it only when a separate check of the captions finds that person on the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.