Key Takeaways

  • Radical Numerics uses mechanistic interpretability from language modeling to inspect what genomic foundation models learn during self-supervised pre-training.
  • Without explicit labels, DNA sequence models distill biological features directly into their internal weights, including GC content, chromatin accessibility, and transcription factor binding motifs.
  • Manual experimentation cannot map the vast combinatorial space of transcription factor interactions; sequence compression learns these distributions directly from data.
  • Radical Numerics is constructing a unified manifold of disease, organizing molecular dysfunction across high-dimensional activation space rather than training separate point models for each disease.

Unsupervised Biology Through Sequence Compression

Biology has long suffered from a cataloging problem. Researchers spend decades testing binding sites and measuring affinities one wet-lab experiment at a time. Eric Nguyen, co-founder of Radical Numerics and creator of the Evo model family, takes a different path. He treats raw biological sequences the same way frontier AI labs treat internet text: as high-dimensional data waiting to be compressed.

“You could think of what the models are doing is compressing a bunch of information it's seen,” Nguyen says. “And so in this compression, it's really basically distilling it down into the key components, key patterns that helps it understand the data or learns the distribution of the data.”

When a model trains on raw genomic sequences, it does not merely memorize nucleotide strings. To predict the next base pair accurately, the network is forced to discover the underlying regulatory logic of the genome. By probing internal representations, Nguyen's team found that models independently learn physical biology. They identify GC content variations, chromatin accessibility patterns, and transcription factor binding motifs with zero supervision.

Mapping the Disease Manifold with Model Activations

The classic approach to computational biology builds narrow classifiers for separate biological questions. You train one model for splicing, another for promoter strength, and a third for pathogenic single-nucleotide variants. Nguyen argues this misses the shared geometric structure of biological dysfunction.

“And in many ways the combinatoral space of learning these transcription factors like what binds and where they bind and what effects it causes is just far too vast to be able to do this in a manual way,” Nguyen explains. “And so we want to take a datadriven approach to to learn some of these motifs.”

Instead of treating each disease variant as an isolated anomaly, Radical Numerics applies mechanistic interpretability techniques borrowed from large language models. They probe model activations across intermediate layers to extract a unified representation space.

“So I think what our first idea for previewing this was and what we're working toward this broadly is this idea of mapping the manifold of disease,” Nguyen notes. “Manifold there's like the sort of representation space of what disease looks like to a model in terms of the output scores or embeddings.”

When you treat disease as a continuous manifold rather than a set of disjointed labels, cross-phenotype relationships become measurable. A mutation in a non-coding region registers as a measurable shift in activation space. While early demonstrations rely on DNA-only architectures like HyenaDNA and Omni, Radical Numerics is expanding this framework to multimodal models that process dozens of biological data types at once.

What to Do With This

Stop building narrow classifiers on hand-engineered biological features. Audit your sequence models this week by attaching linear probes to the intermediate activations of your pre-trained checkpoints, testing specifically for chromatin state and known transcription factor motifs. If your core representations fail to recover these basic regulatory features before fine-tuning, revise your context length and sequence tokenization before training downstream predictors.