Key Takeaways
- DeepMind's Pushmeet Kohli argues AlphaFold 2 cannot be interpretable to people because biological mechanisms exceed the cognitive capacity of the human brain.
- High accuracy is useless without calibration; Kohli notes that a 95 GDT structure prediction model with uncalibrated confidence would trick biologists into wasting an entire year on false positives.
- CZ Biohub's Sal Candido frames protein language models as compressed world models that capture billions of years of evolutionary data rather than simple statistical pattern matchers.
- Model safety requires behavioral characterization, testing operational boundaries and failure cases instead of pretending humans can read neural network attention maps.
The Human Brain Is the Bottleneck
Biologists constantly demand interpretability. They want to know why a model predicts a specific 3D coordinate for an amino acid, hoping for a clean mechanistic explanation they can write on a whiteboard.
Kohli rejects that premise. “Interpretability asks the question interpretable by whom?” Kohli says. “If you're saying interpretable by a human rational system with the cognitive and computational limitations of the human brain, then no, AlphaFold 2 is not interpretable.”
Biology is messy, wide, and nonlinear. When an algorithm folds a complex chain, it factors in thousands of subtle energetic constraints simultaneously. Demanding that a neural network explain itself in terms a human can grasp forces the model to dumb down reality. As Candido explains, these networks build internal world models by compressing evolutionary history: “How does the model do its job? How does it design a protein? Well, it's compressed all the information from evolution into this model.”
If you force an AI to produce simple human heuristics, you strip away the exact high-dimensional patterns that make it work in the first place. Kohli suggests that if we ever decode AlphaFold, it will not come from human inspection. Instead, future frontier models will inspect AlphaFold's activation layers and formulate their own automated theories of its behavior.
Calibration Protects Wet Labs from Ruin
Instead of chasing interpretability, builders must obsess over calibration. In structural biology, AlphaFold uses pLDDT (predicted Local Distance Difference Test) to score its own confidence from 0 to 100 on every residue.
That single metric determines whether an experiment happens or gets scrapped. Kohli emphasizes the danger of deploying high raw performance without calibrated confidence: “Even if it was 95 GDT, but the pLDDT score was completely uncalibrated, who would trust it? It would give good answers, but suddenly tell you here is very confident, and you will be working on it for the next one year and finding out it was completely wrong.”
An experimental team spending fifty thousand dollars and six months synthesizing a protein does not need an essay explaining why an alpha helix formed. They need mathematical certainty that a score of 88 actually maps to an 88 percent probability of physical accuracy. A model that admits ignorance saves millions of dollars; a confident hallucination destroys lab budgets.
Replace Post-Hoc Explanations with Behavioral Testing
Founders building vertical AI tools often burn quarters building feature-importance heatmaps to appease enterprise buyers. That is the wrong path.
Kohli calls for behavioral characterization instead. You do not audit a jet pilot by taking an fMRI of their brain while they fly; you put them in a flight simulator, push them into stalls, and map their failure boundaries. Kohli warns: “We can't just be using these models without having that behavioral characterization, because otherwise it will, rather than being helpful, harm us.”
Systematically stress-test where your model fails. Show your customers exact boundary maps, failure rates on out-of-distribution inputs, and calibrated confidence intervals. That builds earned trust.
What to Do With This
Audit your production model's confidence scores against real-world failure rates this week. Plot your predicted probabilities into ten buckets against observed accuracy on your last 1,000 outputs, and scrap any feature-importance dashboard until your calibration error drops below five percent.