Modal's Dlash: 2-4x LLM Inference Speed, No Quality Loss
Modal's Akshat Bubna reveals Dlash, their block-based speculative decoding, delivers 2-4x LLM inference speed without quality loss. Deploy frontier AI performance now.
40 hours of podcasts, in 5 minutes.
Akshat Bubna, CTO of Modal, discusses the company's evolution from a developer experience platform to one optimized for agent experience and specialized AI workloads. He details Modal's "super cloud" strategy, including elastic sandboxing, advanced autoscaling for inference, and innovative networking like I6PN and RDMA for distributed training. The conversation also covers Modal's contributions to LLM inference with speculative decoding and its capacity management approach in the competitive AI compute market.
Modal's Akshat Bubna reveals Dlash, their block-based speculative decoding, delivers 2-4x LLM inference speed without quality loss. Deploy frontier AI performance now.
Modal's Akshat Bubna explains why building infrastructure for AI agents is mirroring the past decade of dev experience, moving past complex YAML to simple decorators.
Modal CTO Akshat Bubna reveals how their I6PN private IPv6 overlay and RDMA deliver 3 TB/s speeds for serverless distributed AI training and complex agent sandboxes.