Key Takeaways
- In early 2024, OpenRouter built and deleted "Mixture of Models" (MoM) because combining outputs produced worse answers than simply querying the single best frontier model.
- Multi-model fusion failed initially because the top model dramatically outperformed the second and third options on every benchmark.
- Reinforcement learning post-training closed the raw capability gap between frontier labs while creating distinct reasoning styles across models.
- Alex Atallah tested multi-model fusion on complex codebase architecture plans, and every participating model confirmed the fused output beat its own individual response.
Why Mixture of Models Failed in 2024
In early 2024, OpenRouter built a prototype called Mixture of Models (MoM). The idea sounded obvious: run a prompt across several top models, synthesize their outputs into a single response, and give the user the best possible answer.
It did not work.
At the time, one model sat far ahead of the pack. Blending its output with responses from weaker competitors diluted the quality. As Atallah recalled: “In our case the technology was a little too early. In other words the fused result was a little bit worse, sometimes the same as the best model that was being used to fuse, because the best model was so far ahead of options two and three at the time.”
Because the blended answers underperformed the top standalone model, the team scrapped the product. “The form factor was not right and so we would have had to go through a couple more iterations,” Atallah said. “And so we decided to just delete all the code.”
How RL Post-Training Made Councils Work
By 2026, the structural conditions changed. Frontier labs converged on raw capability while diverging in reasoning styles. Reinforcement learning post-training allowed research teams to shape how models explore problem spaces, turning model diversity into an asset rather than a liability.
“Over time, the top three or four LLMs have gotten closer together,” Atallah explained. “Still neurodivergent, but all capable of inserting pretty interesting ideas. RL has basically expanded the surface area of creativity for machine learning researchers within each lab.”
To test whether fusion was ready for real workloads, Atallah ran an experiment on software engineering tasks. He took a complex architecture plan for a code change, distributed it across the top frontier models, and synthesized the collective answers into a single merged plan. Then he ran an automated evaluation step: he fed the fused plan back to each individual model and asked whether the blended architecture was superior to the model's own standalone draft.
Every model agreed that the fused plan was better than what it had produced alone.
When models have comparable baseline capabilities but approach edge cases from distinct angles, fusion stops acting as an averaging filter. It starts acting as a peer review committee that catches blind spots, patches missed edge cases, and surfaces cleaner abstractions.
What to Do With This
Stop routing complex, high-stakes tasks to a single LLM. Set up a pipeline this week for your most difficult prompt (such as a core refactor plan or a strategic document) that queries three distinct frontier models in parallel. Pass all three outputs to a synthesizer prompt with instructions to resolve contradictions and merge the best elements, then evaluate the synthesized response against each individual draft.