Issue No. 37Week ending Sunday, September 13, 2026434 episodes · 1825 articles
The Throughline ↓
The Podcast Summary.

40 hours of podcasts, in 5 minutes.

Guest

Omar Sanseviero

Omar Sanseviero appears in 1 full episode we cover on Latent Space. Below is what each conversation covered, with a key takeaway per article. Every quote in the articles is verbatim and timestamped to the source video.

1 episodecovered
6 articleswith timestamped quotes
AILatent Space

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

Omar Sanseviero from Google DeepMind discusses the release and capabilities of Gemma 4, including its novel E2B architecture for on-device inference, its multimodal and multilingual advancements, and its positioning relative to larger models like Gemini. The conversation also explores broader trends in AI development, such as the evolving role of fine-tuning, the challenges of MOE models, and the increasing convergence of research and engineering practices within Google and the wider AI community.

  • Research now often means "engineering with unknowns": Omar Sanseviero from Google DeepMind notes that many researchers spend their days on "ablations," which means moving pieces around to see what works. He believes this is "much more engineering rather than for like research," unless you are desig… Read →
  • Gemma 4 brings multimodal AI to your pocket: Google DeepMind's Omar Sanseviero confirmed that Gemma 4's smaller models can now process audio, images, and short videos (30-60 seconds) right on a device, a leap for edge computing previously reserved for larger, cloud-based AI. Read →
  • Google's Gemma 4, an on-device model, already matches the state-of-the-art capabilities from 1 to 1.5 years ago for local functions like agentic tasks and conversational AI. Read →
  • Google DeepMind’s Gemma 4 model uses a novel E2B architecture that loads only 2 billion “effective” parameters onto the GPU for inference, even though the model itself contains nearly 5 billion parameters. Read →
  • General fine-tuning for broad conversational changes in large language models is rapidly becoming unnecessary, with models like Google DeepMind's Gemma 4 performing exceptionally well out-of-the-box. Read →
  • Google DeepMind architected Gemma with two distinct models: a 31B dense model for maximum "raw intelligence," and a 27B Mixture-of-Experts (MOE) model optimized for "extremely fast inference" on consumer GPUs. Read →
The Sunday Email

Get next Sunday's issue in your inbox.

40 hours of podcasts, distilled into one 5-minute read. Free, every Sunday morning.

One email a week. Unsubscribe with one click.