Key Takeaways

  • Even with advanced AI drug discovery platforms like Zera Therapeutics’ X-Cell model, the current bottleneck for virtual cell modeling isn't just compute power or algorithms, but the type of data we can actually collect.
  • Ci Chu, from Zera, argues that RNA sequencing, while valuable, only tells part of the story. Proteins are the actual functional units of a cell, and we lack high-throughput methods to measure their abundance, post-translational modifications, localization, and conformational states at the single-cell or spatial level.
  • Bo Wang highlights a critical limitation in current sequencing: you have to destroy a cell to read its state. To build dynamic virtual cells, we need technology that can measure the same cells at multiple time points without killing them.
  • Today’s sophisticated AI models are forced to work with largely static, snapshot data. Unlocking truly dynamic, predictive virtual cell models requires inventing entirely new data generation primitives, not just refining existing ones.

The Protein Blind Spot

For all the excitement around genomic sequencing, the truth is that RNA, while informative, only foreshadows what might happen in a cell. Ci Chu, a guest on Latent Space, cuts straight to the core problem: “RNA is amazing. It foreshadows which proteins are going to get made but protein by and large are the functional units in a cell.” He's not just talking about protein abundance, which is hard enough to measure at scale. He’s pushing for a multi-dimensional view.

Imagine knowing everything about every protein in a single cell: not just how many there are, but their post-translational modifications, their precise location within the cell, and even their current conformational states. Ci Chu sees this as the next frontier. He told the podcast, “If we can measure all of those things, their confirmational states, their modifications, their abundances, their localizations at scale, single cell or even spatially, I think such data sets will be incredibly useful to train the next generation of financial models.” This isn't just more data; it's a qualitatively different kind of data, essential for modeling true biological function.

Cells That Die Before They Talk

Bo Wang brings up an equally pressing, yet often overlooked, data bottleneck: the inherent destructiveness of current cell measurement techniques. Picture trying to understand a complex system by only getting a single snapshot of its state, knowing that taking that snapshot destroys the system itself. That’s essentially what happens with most cell sequencing.

“My hope is I hope to see a breakthrough in sequencing technology,” Wang stated, “Not just the reduced cost but sequencing technology that can sequence the same cells at different time points.” This capability would be a game-changer. Right now, to sequence a cell, you have to kill it. This means all our current data, and thus our virtual cell models, are fundamentally static. We're modeling a fixed moment, not the dynamic process of life. Wang emphasizes the urgency: “Can we have a technology that can measure the cell states at different time points for the same set of cells? I think that will bring a very different dimension to the data set so that we can start to measure the temporal dynamics of cells. So far everything we measure, everything we model is extremely static.” Unlocking temporal dynamics is the key to truly simulating a 'virtual cell' as a living, evolving entity.

What to Do With This

If you're building AI in biology, stop chasing marginal gains on existing data. Instead, deploy your smartest engineers to solve the data generation challenge. Specifically, focus on non-destructive, longitudinal single-cell measurement or high-throughput, multi-dimensional protein profiling. The first team to crack either of these will unlock a multi-billion dollar market and redefine what's possible in drug discovery and cell engineering.