Key Takeaways
- Sal Candido warns that scaling laws do not appear automatically in biology; discovering the exact regimes where data volume predictably drives performance is an open empirical problem.
- Training protein language models on messy metagenomic fragments, where sequences do not represent complete proteins, still improves real protein design performance.
- Pushmeet Kohli argues that Rich Sutton's Bitter Lesson warns against dogmatic identity politics between modelers and wet-lab experimentalists.
- DeepMind tested single-cell transcriptomics data in CELLxGENE to evaluate building a "virtual cell" and discovered the data volume was too sparse to scale.
Biology Does Not Hand You Scaling Laws for Free
Silicon Valley treats Rich Sutton's Bitter Lesson like a universal law: stop building clever heuristics, feed more compute and raw tokens into a transformer, and watch loss curves decline.
In biology, that assumption breaks quickly.
Sal Candido from CZ Biohub pointed out that machine learning engineers confuse the existence of scaling laws with their discovery. “I think one misconception of scaling laws is that scaling laws are everywhere and they always exist,” Candido explained. “I think a lot of the work is actually finding that scaling law.”
In natural language, internet scrape data offers a direct path to scaling. In biological systems, compute only helps if you find a data modality that contains the underlying generative grammar of life. When you find that grammar, the results surprise even the engineers who build them.
Candido described how CZ Biohub trains protein language models on messy metagenomic fragments: “We train on metagenomic sequences which are not the highest quality data. In fact, much of that data I can guarantee you isn't even a real whole protein. And yet that makes the performance of the model go up for designing real proteins that work, for understanding proteins that we know are things.”
Even partial, broken sequence fragments from uncultivated microbes provide enough signal about evolutionary constraints to teach a neural network how to design viable molecules. The scale works because evolutionary sequences carry a coherent statistical structure, even when the laboratory capture method is noisy.
The Virtual Cell Wall and the Model Dogma
Pushmeet Kohli, head of AI for science at Google DeepMind, takes the Bitter Lesson argument a step further. To Kohli, Sutton's warning was never an order to abandon wet-lab experimenters in favor of pure GPU clusters. Sutton was attacking dogma.
“Sometimes when we are looking at problems we think about solutions in a very religious way,” Kohli said. “'I'm a modeler' or 'I'm a data generation person.' And I think that is the bitter lesson, that if you approach the problem with that mindset, you might not succeed. The problem comes first and you should be flexible in terms of your solution space.”
When DeepMind tried to extend its post-AlphaFold work into single-cell biology, the limits of pure modeling became obvious immediately. The team investigated whether massive single-cell RNA collections could anchor an end-to-end model of cellular behavior.
“When you're approaching say cell genomics, then we took the same approach and said, well, what can you do with cell by gene?” Kohli noted, referring to CZ CELLxGENE datasets. “And there it was very clear after a bunch of work that the data was not there yet to be able to build, to go after that grand ambition of building the virtual cell.”
If the data foundation does not exist, throwing larger clusters of GPUs at the problem produces zero returns. You cannot compute your way out of missing measurements. In those regimes, builders must step back from training runs, enter the wet lab, and generate the physical assays required to establish a real signal.
What to Do With This
Audit your technical roadmap tomorrow morning. If your team is tweaking model architectures on an existing dataset without seeing measurable performance jumps, stop optimizing. Run a downsampling test across 25%, 50%, and 100% of your training data: if the performance curve is flat, you do not have a modeling bottleneck, you have an unscalable data regime that requires building new assays.