Issue No. 21Week ending Sunday, May 24, 2026436 episodes · 1837 articles
The Throughline ↓
The Podcast Summary.

40 hours of podcasts, in 5 minutes.

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

With swyx, Alessio Fanelli, Omar Sanseviero · Sunday, May 24, 2026

Omar Sanseviero from Google DeepMind discusses the release and capabilities of Gemma 4, including its novel E2B architecture for on-device inference, its multimodal and multilingual advancements, and its positioning relative to larger models like Gemini. The conversation also explores broader trends in AI development, such as the evolving role of fine-tuning, the challenges of MOE models, and the increasing convergence of research and engineering practices within Google and the wider AI community.

Key takeaways

  • Research now often means "engineering with unknowns": Omar Sanseviero from Google DeepMind notes that many researchers spend their days on "ablations," which means moving pieces around to see what works. He believes this is "much more engineering rather than for like research," unless you are designing new fundamental architectures. Read more →
  • Gemma 4 brings multimodal AI to your pocket: Google DeepMind's Omar Sanseviero confirmed that Gemma 4's smaller models can now process audio, images, and short videos (30-60 seconds) right on a device, a leap for edge computing previously reserved for larger, cloud-based AI. Read more →
  • Google's Gemma 4, an on-device model, already matches the state-of-the-art capabilities from 1 to 1.5 years ago for local functions like agentic tasks and conversational AI. Read more →
  • Google DeepMind’s Gemma 4 model uses a novel E2B architecture that loads only 2 billion “effective” parameters onto the GPU for inference, even though the model itself contains nearly 5 billion parameters. Read more →
  • General fine-tuning for broad conversational changes in large language models is rapidly becoming unnecessary, with models like Google DeepMind's Gemma 4 performing exceptionally well out-of-the-box. Read more →
  • Google DeepMind architected Gemma with two distinct models: a 31B dense model for maximum "raw intelligence," and a 27B Mixture-of-Experts (MOE) model optimized for "extremely fast inference" on consumer GPUs. Read more →

6 articles from this episode

AI Agents: When Research Becomes Engineering

Omar Sanseviero of Google DeepMind reveals how AI agents are turning research into automated engineering, freeing human experts to find entirely new discoveries.

Read article →

More Latent Space episodes

Every Latent Space episode we cover →

The Sunday Email

Get next Sunday's issue in your inbox.

40 hours of podcasts, distilled into one 5-minute read. Free, every Sunday morning.

One email a week. Unsubscribe with one click.