Key Takeaways

  • ReflectionAI trained Beam, a 500-billion-parameter model with 23 billion active parameters, spending more FLOPs on reinforcement learning than on pre-training.
  • Pre-training Beam took 6,000 GB300 GPUs for several weeks, while reinforcement learning consumed over 10,000 GB300s for four solid weeks.
  • Beam runs three to four times more efficiently than models in its capability class, reaching up to a 10x efficiency advantage against larger models.
  • Post-training reinforcement learning curves show no performance wall yet; progress now depends on capital allocation and synthetic data generation rather than raw web text.

The Compute Inversion: Why RL Outspent Pre-Training

For years, the standard playbook in artificial intelligence was simple. You poured 95 percent of your budget into pre-training on public text, then ran a quick pass of supervised fine-tuning and basic reinforcement learning at the finish line.

ReflectionAI flipped that budget on its head to build Beam.

Beam is a 500-billion-parameter reasoning model with 23 billion active parameters. To get it off the ground, the team needed 6,000 Nvidia GB300 GPUs running for several weeks. With recent cluster optimizations, Misha Laskin notes they can now hit that same pre-training milestone in about 12 days.

The real compute bill came after the base model was finished.

“Reinforcement learning was a little over 10,000 GB300s for four weeks,” Laskin said. “So, actually there are more flops spent on reinforcement learning.”

That run marks the largest reinforcement learning effort in open-weight history. It also signals a permanent shift in how teams produce intelligence. Pre-training gives a model its world facts and language grammar, but raw text runs out. Reinforcement learning trains the model to think through problems, test hypotheses, and correct mistakes.

Even better, the returns have not flattened. “And if you look at our plots, they just keep going up and it's just a matter of compute basically,” Laskin said. “We as a field are there where you have these RL systems that are very general that don't really stop improving.” The bottleneck is no longer scientific theory. It is your ability to purchase compute and generate high-grade synthetic data.

The Coupling Trap: You Still Cannot Skip the Base

When reasoning models first started showing massive jumps via reinforcement learning, some researchers wondered if pre-training would disappear entirely. If an RL loop can teach a model to solve math and code, why spend millions scraping the open internet?

Laskin warns that this shortcut fails in practice. You cannot take an untrained network and point an RL reward signal at it.

“It turned out you actually do need to pre-train your model in order to make reinforcement learning work very well at scale,” Laskin explained. “The things are just so tightly coupled that you do need to do both.”

Pre-training builds the latent representation of concepts. RL acts like an intense coaching program that teaches the model how to traverse those concepts without getting confused. If the foundation is missing, the coach has nothing to work with.

Because Beam activates only 23 billion of its 500 billion parameters per forward pass, ReflectionAI gets the reasoning depth of a giant model with the inference speed of a small one. “Beam tends to be three to four times more efficient than models of the same capability class,” Laskin noted, adding that against larger architectures, “the efficiency gains are then end up being something like 10x.”

What to Do With This

Audit your model roadmap before you raise your next round. If your technical plan still allocates 80 percent of compute to collecting raw domain text for base pre-training, rebalance your budget toward synthetic problem generation and RL verification environments. Build verifiers that can automatically score model outputs in your vertical tomorrow, because that is where reasoning actually scales.