5 quotes from 1 episode on Latent Space, each with a timestamped link to the source.
5 quotes1 episode
The short version
Akshat Bubna explains that AI infrastructure teams now build tools for agents and optimize inference speed. Batching the verification of tokens predicted by a small draft model yields a 2-4x speedup in compute efficiency.
Most interesting insights
Speculative decoding relies on a small draft model predicting tokens before a large model verifies the output.
“…speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model and then you have the bigger model verify all of this.”
Batching draft model verification increases efficiency
A smaller draft model predicts tokens ahead of time. Batching the verification step for a large model improves compute efficiency and generates a 2-4x speedup.
“…if you can batch the verification of the draft model then you're much more efficient using compute and it's faster.”
Infrastructure teams optimize for agent experience
Software development kits now target AI agents as the primary users. Akshat Bubna shifted a team focus to agent experience to make infrastructure accessible through simple code decorators.
“We've actually changed our SDK team to think about agent experience instead of developer experience…”
Modal's SDK team pivoted from Developer Experience (DX) to Agent Experience (AX), recognizing that AI agents are now the primary consumers of infrastructure.
Just as developers hated complex Kubernetes YAML, agents shouldn't have to either. Akshat Bubna, Modal's CTO, argues infrastructure should be accessible via simple code decorators.
Modal has open-sourced Dlash, a block-based speculative decoding technique designed to accelerate LLM inference without compromising output quality.
Dlash achieves a 2-4x speedup by employing a smaller "draft" model to predict tokens ahead, allowing the larger model to verify predictions in batches, efficiently using compute.
How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.