Baseten's Long-Context LLMs: Cache, Disaggregate, Speculate
Founders, learn how Baseten efficiently handles 200,000-token LLM queries with cache-aware routing, prefill/decode disaggregation, and speculative decoding.
40 hours of podcasts, in 5 minutes.
This episode of Latent Space features Philip Kiely and Ali Taha from Baseten, discussing advanced inference engineering for AI models. They delve into strategies for optimizing model performance, reliability, and cost, particularly for large language models, and explore the future of hardware-software co-design. Key topics include managing long context queries, the complex process of supporting new model releases, novel quantization techniques, and the challenges of video generation.
Founders, learn how Baseten efficiently handles 200,000-token LLM queries with cache-aware routing, prefill/decode disaggregation, and speculative decoding.
Baseten's Ali Taha reveals how mathematically proven quantization can increase LLM throughput by 20% and improve fidelity. Stop trading quality for speed.
Under Nvidia's Rubin, AI inference engineering demands less low-level kernel tuning and more systems orchestration. Kiely and Taha debate the future.
Ali Taha from Baseten reveals how their GLM 5.2 model autonomously analyzes, writes, and implements its own optimized GPU kernels for faster inference.
Baseten's Ali Taha explains why diffusion models hit a quadratic bottleneck for long video and why auto-regressive architectures are the clear future.