Key Takeaways

  • Protein-coding regions make up only 1.5% to 2% of the human genome, leaving 98% of genetic variation in poorly understood non-coding regulatory sequences.
  • Most clinical variant predictors rely on specialized supervised heads, which fail when applied outside narrow protein-coding benchmarks.
  • Radical Numerics built Omni to predict pathogenic mutations directly out of the box by pairing unsupervised likelihood scoring with structured mid-training prompts.
  • Omni matches or beats specialized bioinformatic models on standard clinical benchmarks like ClinVar and TraitGym without requiring custom downstream fine-tuning.

The 98% Blind Spot in Modern Genomics

Biotech software spent decades focusing almost entirely on the 2% of human DNA that codes directly for proteins. It made sense early on: amino acid chains fold into physical shapes, making errors easier to spot under a microscope or inside a structural prediction model. But that leaves the other 98% of the sequence, the regulatory switches that control when and where proteins get built, as an unsolved problem.

Most disease-causing mutations lurk inside these non-coding stretches. When clinicians run genetic panels on sick patients, they run into a wall of unknown variants. Eric Nguyen puts the scale of the challenge in plain numbers: “Because the combinations of these changes in the genome over three billion letters is so vast that for clinicians and scientists, we actually know only a very small portion of which variants are causal to disease or not.”

Traditional bioinformatic tools rely on supervised classifiers trained on labeled datasets. That method works well when you have millions of verified examples from protein databases. In non-coding DNA, where wet-lab labels are scarce and noisy, supervised classifiers collapse. Nguyen notes: “In many cases these non-coding regions, these regulatory regions variants there, mutations there are much harder to predict if they cause disease or not. And I think what's exciting about the new generation of models we're building with Omni is that that's where we shine.”

Unsupervised Likelihood Beats Task-Specific Heads

The standard machine learning playbook in biology involves taking a pre-trained base model, chopping off its final layers, and training a regression head on a specific benchmark like ClinVar. That approach locks the architecture into narrow tasks and makes multi-task generalization impossible.

Omni takes a different route. Instead of post-hoc supervised heads, the model relies on zero-shot likelihood scoring across massive raw genomic sequences. During pre-training and mid-training, the model learns the statistical distribution of evolutionary variation across species. When presented with a candidate mutation, it evaluates how sharply that mutation drops the sequence likelihood score relative to the wild-type reference.

To make this work across disparate biological tasks, Radical Numerics incorporated structured prompt formatting directly into the mid-training phase. Nguyen explains the architectural benefit: “We've done it so that the model is flexible during mid training to be trained on many tasks at once. And so the fine-tuning, that last step of training the embeddings, doesn't have to be done. It's out of the box at that point.”

By treating DNA as a continuous sequence without task-specific patches, Omni handles long-range regulatory dependencies that traditional windowed tools clip or miss entirely.

What to Do With This

Stop adding custom regression heads on top of static embeddings for your biological data pipelines. Run a zero-shot log-likelihood test on your raw sequence model against a small validation slice of non-coding variants this week. If the base model's raw probability distribution ranks pathogenic mutations better than your fine-tuned head, strip the downstream classifier entirely.