⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Omar Sanseviero from Google DeepMind discusses the release and capabilities of Gemma 4, including its novel E2B architecture for on-device inference, its multimodal and multilingual advancements, and its positioning relative to larger models like Gemini. The conversation also explores broader trends in AI development, such as the evolving role of fine-tuning, the challenges of MOE models, and the increasing convergence of research and engineering practices within Google and the wider AI community.
- Research now often means "engineering with unknowns": Omar Sanseviero from Google DeepMind notes that many researchers spend their days on "ablations," which means moving pieces around to see what works. He believes this is "much more engineering rather than for like research," unless you are desig… Read →
- Gemma 4 brings multimodal AI to your pocket: Google DeepMind's Omar Sanseviero confirmed that Gemma 4's smaller models can now process audio, images, and short videos (30-60 seconds) right on a device, a leap for edge computing previously reserved for larger, cloud-based AI. Read →
- Google's Gemma 4, an on-device model, already matches the state-of-the-art capabilities from 1 to 1.5 years ago for local functions like agentic tasks and conversational AI. Read →
- Google DeepMind’s Gemma 4 model uses a novel E2B architecture that loads only 2 billion “effective” parameters onto the GPU for inference, even though the model itself contains nearly 5 billion parameters. Read →
- General fine-tuning for broad conversational changes in large language models is rapidly becoming unnecessary, with models like Google DeepMind's Gemma 4 performing exceptionally well out-of-the-box. Read →
- Google DeepMind architected Gemma with two distinct models: a 31B dense model for maximum "raw intelligence," and a 27B Mixture-of-Experts (MOE) model optimized for "extremely fast inference" on consumer GPUs. Read →