Issue No. 40Week ending Sunday, October 4, 2026485 episodes · 2075 articles
The Throughline ↓
The Podcast Summary.

10+ hours of podcasts, in 5 minutes.

AI infrastructure and compute

Akshat Bubna on AI infrastructure and compute

5 quotes from 1 episode on Latent Space, each with a timestamped link to the source.

5 quotes1 episode

The short version

Akshat Bubna explains that AI infrastructure teams now build tools for agents and optimize inference speed. Batching the verification of tokens predicted by a small draft model yields a 2-4x speedup in compute efficiency.

Most interesting insights

Speculative decoding relies on a small draft model predicting tokens before a large model verifies the output.

“…speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model and then you have the bigger model verify all of this.”

Akshat Bubna, Latent Space · July 2026 · Watch at 18:05 ↗

From Modal's Dlash: 2-4x LLM Inference Speed, No Quality Loss

Dashboard quality dictates how well teams observe and monitor AI systems.

“…one thing we actually still see is really important is observability. How good is your dashboard?”

Akshat Bubna, Latent Space · July 2026 · Watch at 6:59 ↗

From Modal CTO: Agent Experience Is the New Developer Experience

Human workers evaluate system performance and make judgment calls on the collected data.

“You still need humans to go interpret what's going on and make judgment calls and whatnot…”

Akshat Bubna, Latent Space · July 2026 · Watch at 7:11 ↗

From Modal CTO: Agent Experience Is the New Developer Experience

Top talking points

  1. Batching draft model verification increases efficiency

    A smaller draft model predicts tokens ahead of time. Batching the verification step for a large model improves compute efficiency and generates a 2-4x speedup.

    “…if you can batch the verification of the draft model then you're much more efficient using compute and it's faster.”

    Akshat Bubna, Latent Space · July 2026 · Watch at 18:24 ↗

    From Modal's Dlash: 2-4x LLM Inference Speed, No Quality Loss

  2. Infrastructure teams optimize for agent experience

    Software development kits now target AI agents as the primary users. Akshat Bubna shifted a team focus to agent experience to make infrastructure accessible through simple code decorators.

    “We've actually changed our SDK team to think about agent experience instead of developer experience…”

    Akshat Bubna, Latent Space · July 2026 · Watch at 6:05 ↗

    From Modal CTO: Agent Experience Is the New Developer Experience

Key takeaways from these write-ups

Modal CTO: Agent Experience Is the New Developer Experience

  • Modal's SDK team pivoted from Developer Experience (DX) to Agent Experience (AX), recognizing that AI agents are now the primary consumers of infrastructure.
  • Just as developers hated complex Kubernetes YAML, agents shouldn't have to either. Akshat Bubna, Modal's CTO, argues infrastructure should be accessible via simple code decorators.

Modal's Dlash: 2-4x LLM Inference Speed, No Quality Loss

  • Modal has open-sourced Dlash, a block-based speculative decoding technique designed to accelerate LLM inference without compromising output quality.
  • Dlash achieves a 2-4x speedup by employing a smaller "draft" model to predict tokens ahead, allowing the larger model to verify predictions in batches, efficiently using compute.

How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.

More

The Sunday Email

Get next Sunday's issue in your inbox.

10+ hours of podcasts, distilled into one 5-minute read. Free, every Sunday morning.

Newsletters

One email a week. Unsubscribe with one click.