5 quotes from 1 episode on Latent Space, each with a timestamped link to the source.
5 quotes1 episode
The short version
Philip Kiely details how engineers redesign AI infrastructure to match the exact workload shapes of large language models. Systems now divide query prefill and decoding across separate GPU groups to handle 200,000-token inputs.
Most interesting insights
The Rubin chip stands as the first architecture designed specifically around the workload shapes of large language models.
“Rubin's honestly the first chip that was fully built in that world. And so, you can see a lot of the understanding of the shape of the workload that this chip's going to be asked to do in the way it's designed.”
Leading agent builders will run production loops that learn directly from active inference data within a few months to a couple of years.
“I think within a few months to a couple years like a lot of leading agent builders are going to have these loops like really up and running in production where you are doing influence learning from the influence.”
Philip Kiely notes that systems check if parts of a 200,000-token query appeared before. Networks then route the query to available prefill workers that already hold the cached data to skip redundant processing.
“With a long query specifically, the first thing that I'm going to ask is, have you sent me this query before, or at least part of it?”
“…we want to send this one to something with number one available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of this 200,000 tokens.”
One group of GPUs processes the initial query and cache to generate the first token. A completely separate set of GPUs then takes over to run the decode step.
“…we've at least on certain models disaggregated prefill and decode. You're going to have one set of GPUs that's solely going to process the input query, that KV cache, and get you your first token. And then that's going to be passed over to a separate set of GPUs which is going to run decode.”
To handle massive 200,000-token LLM queries cost-effectively, Baseten first checks if parts of the input have been seen before, using cache-aware routing to skip expensive re-computations.
They split the LLM inference pipeline into two distinct stages—prefill and decode—running them on separate sets of GPUs to optimize resource utilization and throughput for mixed workloads.
Nvidia's upcoming Rubin architecture signals a new era for AI inference, specifically designed with a deep understanding of large language model workloads from its inception.
The core debate centers on whether AI inference engineering will demand more systems-level thinking (like CPU-GPU interconnects and KV cache management) or if it will become purely an infrastructure orchestration problem.
AI models are now getting good at autonomously optimizing their own underlying infrastructure. Baseten’s GLM 5.2 model proves this by rewriting its own GPU kernels.
This creates a continuous improvement loop: live inference data directly feeds into the model’s ability to refine its performance, not just its outputs.
How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.