Issue No. 40Week ending Sunday, October 4, 2026485 episodes · 2075 articles
The Throughline ↓
The Podcast Summary.

10+ hours of podcasts, in 5 minutes.

AI infrastructure and compute

Philip Kiely on AI infrastructure and compute

5 quotes from 1 episode on Latent Space, each with a timestamped link to the source.

5 quotes1 episode

The short version

Philip Kiely details how engineers redesign AI infrastructure to match the exact workload shapes of large language models. Systems now divide query prefill and decoding across separate GPU groups to handle 200,000-token inputs.

Most interesting insights

The Rubin chip stands as the first architecture designed specifically around the workload shapes of large language models.

“Rubin's honestly the first chip that was fully built in that world. And so, you can see a lot of the understanding of the shape of the workload that this chip's going to be asked to do in the way it's designed.”

Philip Kiely, Latent Space · August 2026 · Watch at 1:06:13 ↗

From Rubin Era: AI Inference Shifts From Kernels to Orchestration

Leading agent builders will run production loops that learn directly from active inference data within a few months to a couple of years.

“I think within a few months to a couple years like a lot of leading agent builders are going to have these loops like really up and running in production where you are doing influence learning from the influence.”

Philip Kiely, Latent Space · August 2026 · Watch at 1:31:14 ↗

From Baseten's GLM 5.2 Rewrites Its Own GPU Kernels

Top talking points

  1. Caching inputs speeds up long queries

    Philip Kiely notes that systems check if parts of a 200,000-token query appeared before. Networks then route the query to available prefill workers that already hold the cached data to skip redundant processing.

    “With a long query specifically, the first thing that I'm going to ask is, have you sent me this query before, or at least part of it?”

    Philip Kiely, Latent Space · August 2026 · Watch at 2:47 ↗

    From Baseten's Long-Context LLMs: Cache, Disaggregate, Speculate

    “…we want to send this one to something with number one available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of this 200,000 tokens.”

    Philip Kiely, Latent Space · August 2026 · Watch at 3:12 ↗

    From Baseten's Long-Context LLMs: Cache, Disaggregate, Speculate

  2. Prefill and decode require separate hardware

    One group of GPUs processes the initial query and cache to generate the first token. A completely separate set of GPUs then takes over to run the decode step.

    “…we've at least on certain models disaggregated prefill and decode. You're going to have one set of GPUs that's solely going to process the input query, that KV cache, and get you your first token. And then that's going to be passed over to a separate set of GPUs which is going to run decode.”

    Philip Kiely, Latent Space · August 2026 · Watch at 3:38 ↗

    From Baseten's Long-Context LLMs: Cache, Disaggregate, Speculate

Key takeaways from these write-ups

Baseten's Long-Context LLMs: Cache, Disaggregate, Speculate

  • To handle massive 200,000-token LLM queries cost-effectively, Baseten first checks if parts of the input have been seen before, using cache-aware routing to skip expensive re-computations.
  • They split the LLM inference pipeline into two distinct stages—prefill and decode—running them on separate sets of GPUs to optimize resource utilization and throughput for mixed workloads.

Rubin Era: AI Inference Shifts From Kernels to Orchestration

  • Nvidia's upcoming Rubin architecture signals a new era for AI inference, specifically designed with a deep understanding of large language model workloads from its inception.
  • The core debate centers on whether AI inference engineering will demand more systems-level thinking (like CPU-GPU interconnects and KV cache management) or if it will become purely an infrastructure orchestration problem.

Baseten's GLM 5.2 Rewrites Its Own GPU Kernels

  • AI models are now getting good at autonomously optimizing their own underlying infrastructure. Baseten’s GLM 5.2 model proves this by rewriting its own GPU kernels.
  • This creates a continuous improvement loop: live inference data directly feeds into the model’s ability to refine its performance, not just its outputs.

How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.

More

The Sunday Email

Get next Sunday's issue in your inbox.

10+ hours of podcasts, distilled into one 5-minute read. Free, every Sunday morning.

Newsletters

One email a week. Unsubscribe with one click.