Key Takeaways

  • Academic journals create massive data bias because researchers only publish successful crystal syntheses and throw away failed reactions.
  • Machine learning models trained only on polished papers cannot troubleshoot physical lab work because they never see negative samples.
  • Liam Fedus and Ekin Dogus Cubuk build automated labs that log the full lineage of an experiment: researcher chats, raw code, computational runs, and broken attempts.
  • Feeding end-to-end failure loops into models prevents reinforcement learning systems from gaming benchmarks through memorized answers.
  • Training on the mess of experimentation teaches foundation models actual process engineering instead of static textbook answers.

The Negative Data Deficit in Modern Science

Every published chemistry paper tells a clean story: a team chose a hypothesis, ran a reaction, and produced a pristine crystal. Real wet lab work looks nothing like that. It is a grind of wrong temperatures, contaminated solvents, clogged tubing, and days of null results.

Because scientific journals reject papers about reactions that failed, academia buries its mistakes. That selective memory cripples artificial intelligence.

As Cubuk explains: “if you don't have negative samples you can't really train if everything is positive and this is a particularly bad problem in material science because people usually publish crystals they could synthesize but they usually don't publish if they fail to synthesize a crystal sometimes they might.”

When a model ingests millions of papers that only report wins, it develops an unrealistic view of the physical world. It cannot classify reaction feasibility accurately because it never sees the boundary lines where physics says no. Simulators like Density Functional Theory offer approximations, but digital math cannot replace what happens inside an actual furnace or deposition chamber.

Training on the Process, Not the Trophy

Periodic Labs takes the opposite approach. Instead of scraping journals for finished recipes, the team records every breadcrumb of the search process inside automated physical laboratories. They capture the initial conversation between researchers, the code written to operate the hardware, the intermediate diagnostic measurements, and the physical failures.

Fedus describes the value of this complete record: “This type of data basically doesn't exist anywhere else. And we spend so much of our time getting the full lineage of the scientific process into the model. And so like tracking all this data like the conversations, the intuitions, what was executed in the lab, where were the computations run, where was the code written, stitching all this together is is so valuable.”

When you show a model thirty failed attempts followed by the one tweak that cracked the problem, you give it something rare: “I think there's also like another valid thing too of when you have this sort of string of negative results and then finally through process iteration you're able to get to that positive result. It's a really interesting set of I'll call it like process engineering type data,” Fedus notes. “and I think like you know the overall goal is rather than training on the final output of science, you're training on the process of doing science.”

This provenance record fixes another headache in AI development: reward hacking. When models run reinforcement learning loops on open benchmarks, they often cheat by regurgitating memorized literature answers. When they must reason through a complete trail of lineage data, they have to learn real causal physics.

What to Do With This

Audit your internal data collection pipeline tomorrow morning. If your team only logs won deals, merged pull requests, or successful production runs, you are starving your fine-tuning pipeline of the contrast data it needs. Start saving the discarded drafts, rejected code revisions, and failed operational workflows to build models that actually understand how your systems break.