Key Takeaways

  • Pre-training on raw DNA created models that could read and write biology, but raw base models consistently lagged behind narrow specialized tools on human disease-variant prediction.
  • HyenaDNA proved that long-context convolutions could scale genomic models to one million tokens, unlocking sequence lengths that standard attention mechanisms could not process.
  • Evo marked the transition to generative genomics, producing the first complete bacteriophage genome designed from scratch with machine learning.
  • Omni closes the performance gap against specialized tools by introducing mid-training alignment, mechanistic interpretability, and in silico biological chain-of-thought.
  • Frontier biology labs face an unavoidable dual mandate: generating biological sequences requires building AI-native functional defenses against weaponized sequence generation.

The Raw DNA Trap

Most biological foundation models stop at base pre-training. They ingest massive archives of raw sequence data, predict the next base pair, and hope downstream reasoning emerges on its own. In practice, it rarely does.

Eric Nguyen noticed this wall early. While building HyenaDNA, his team solved the context length barrier. Standard attention mechanisms choked on long genomic sequences. The Hyena operator used long convolutions instead, opening up sequence windows up to one million tokens, which stood as the largest context window for a language model at the time.

Scaling context let models see whole genomic structures. That unlocked Evo, which moved the field from passive classification to generative genomics. Researchers used Evo to design a functional bacteriophage genome from scratch. But when applied to precise human genetics, such as predicting whether a specific mutation causes a disease, raw generative models still fell short of simple, specialized baselines.

Nguyen explains the disconnect plainly: “The way we think about language models in bio and genomes in particular is that mostly you've only seen base models trained. So they're pre-trained, but they're basically unaligned.”

Aligning Biology with In Silico Reasoning

In natural language processing, nobody deploys raw base models directly to end users. You apply instruction tuning, reinforcement learning, and structured reasoning techniques.

Nguyen and his team at Radical Numerics applied this playbook to genomics with Omni. Instead of simply predicting the next nucleotide, Omni introduces mid-training alignment and biological chain-of-thought reasoning. The model reasons through intermediate biochemical steps in silico before predicting functional outcomes.

This intermediate reasoning step changes performance entirely. When a model simulates molecular interactions and explains its step-by-step logic, its predictions on disease variants surpass narrow legacy classifiers. It also provides mechanistic interpretability. Biologists no longer have to trust an unexplainable prediction; they can inspect the structural and regulatory logic the model used.

The Functional Defense Mandate

When AI can write viable viral genomes from scratch, defensive capability cannot rely on static databases. Legacy biosecurity depends on watchlists of known dangerous pathogens. If a bad actor orders smallpox DNA from a synthesis provider, the sequence gets flagged against a list.

Generative models break that static filter. A model can generate functional sequences from scratch that execute toxic mechanisms without matching any known watchlist entry.

Nguyen argues that any team building generative biology models must operate under a dual mandate. Building the engine requires building the defense. Defensive tools must be functional rather than signature-based. They must use foundation models to screen sequences based on what they will actually do inside a living cell, not whether their text matches an existing database.

What to Do With This

Audit your domain-specific AI stack this week. If you rely strictly on large base models with zero-shot prompting to solve complex technical tasks, benchmark their accuracy against your best narrow heuristics. If the base model loses, design a mid-training or post-training pipeline that forces the model to generate intermediate reasoning traces before outputting its final prediction.