Key Takeaways

  • DeepMind designed AlphaFold 2 by hardcoding biophysical rules into the network rather than relying on pure compute scaling.
  • Pushmeet Kohli argues that encoding how amino acid residues interact gives biology models an unfair data-efficiency advantage.
  • Sal Candido warns that wrong inductive biases cap model performance as biological training sets expand.
  • Scaling biological AI requires data coverage across distinct molecular states rather than simply collecting more samples of known structures.

The Disagreement

Brandon Anderson posed a central question to AI biology researchers: should teams build bespoke architectures loaded with physical assumptions, or should they feed raw data into standard transformers?

Pushmeet Kohli defends the bespoke approach that powered AlphaFold 2. In biological systems, waiting for a model to infer basic physics from sparse experimental readings wastes millions of dollars and years of training time. Kohli explains that DeepMind acted with specific intent:

“There was a vision behind it that all this scientific intuition that came from biophysics and biochemistry, that those interactions that amino acid residues are not just doing their own thing, they are being influenced by other residues. So let's bake that in. If you have learned something from the scientific community, use that information and try to encourage the model and give it that unfair advantage.”

Sal Candido pushes back on treating those hardcoded rules as permanent solutions. While Candido agrees that domain biases rescue models when data is scarce, he points out the trap that catches teams as their databases expand:

“As you get more and more data, you see that sometimes the model can find things that you didn't necessarily know about, and sometimes that inductive bias, if it wasn't exactly correct, can hold you back.”

Candido notes that while biology is not in a post-transformer world, teams must modify those core architectures to make them fit for purpose at scale rather than freezing scientific dogmas into silicon.

Who's Right (and When They're Wrong)

Kohli is right whenever experimental data is expensive, rare, or physically constrained. Structural biology does not have the billions of cheap tokens that text models enjoy. The Protein Data Bank contains fewer than three hundred thousand structures. If you strip away structural priors and ask a blank transformer to solve protein folding from scratch on that limited dataset, it fails. Handcrafting geometric constraints lets the model skip thousands of unphysical dead ends.

Candido is right once experimental pipelines start producing billions of readouts. Human understanding of biology contains errors, omissions, and historical biases. If your architecture forces a model to obey a 1970s textbook assumption about molecular kinetics, the model cannot discover interactions that violate that textbook. When data collection scales past the point of hand-curated samples, hardcoded assumptions stop acting like training wheels and start acting like cages.

Data strategy breaks the tie. Kohli emphasizes that raw volume does not solve the problem on its own. Replicating the same structural data gives diminishing returns. Progress requires high coverage across unexplored chemical configurations. If your dataset lacks structural diversity, you need Kohli's inductive biases to stay grounded in reality. If you have built high-throughput wet labs that generate diverse, high-coverage assays, Candido's warning applies: loosen your architectural constraints and let the transformer detect patterns human biophysicists missed.

What to Do With This

Audit the core assumptions hardcoded into your model pipeline this week. List the three strongest physical or domain constraints your engineers baked into your architecture, loss function, or feature preprocessing. For each constraint, check your training data volume: if your dataset has grown more than ten times since you wrote that rule, run an ablation experiment removing that constraint. If removing the constraint improves test set accuracy on rare edge cases, delete the rule and let the architecture scale.