Issue No. 32Week ending Sunday, August 9, 2026539 episodes · 2375 articles
The Throughline ↓
The Podcast Summary.

10+ hours of podcasts, in 5 minutes.

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten

With swyx, Philip Kiely, Ali Taha · Sunday, August 9, 2026

This episode of Latent Space features Philip Kiely and Ali Taha from Baseten, discussing advanced inference engineering for AI models. They delve into strategies for optimizing model performance, reliability, and cost, particularly for large language models, and explore the future of hardware-software co-design. Key topics include managing long context queries, the complex process of supporting new model releases, novel quantization techniques, and the challenges of video generation.

Key takeaways

  • To handle massive 200,000-token LLM queries cost-effectively, Baseten first checks if parts of the input have been seen before, using cache-aware routing to skip expensive re-computations. Read more →
  • Conventional wisdom on quantization is dead: Baseten's research shows you can quantize more and still boost LLM quality. Read more →
  • Nvidia's upcoming Rubin architecture signals a new era for AI inference, specifically designed with a deep understanding of large language model workloads from its inception. Read more →
  • AI models are now getting good at autonomously optimizing their own underlying infrastructure. Baseten’s GLM 5.2 model proves this by rewriting its own GPU kernels. Read more →
  • Current open-source video generation models, like 1.2.2, show a "night and day" quality gap compared to cutting-edge models like Kling or Veo, especially for long-form content. Read more →

5 articles from this episode

More Latent Space episodes

Every Latent Space episode we cover →

The Sunday Email

Get next Sunday's issue in your inbox.

10+ hours of podcasts, distilled into one 5-minute read. Free, every Sunday.

Newsletters

For now, every subscriber gets both newsletters. No spam. Unsubscribe with one click.