Key Takeaways

  • Snorkel AI closed a new $350 million valuation round based on a clear thesis: raw data volume no longer improves frontier AI models.
  • Ratner started researching programmatic data labeling a decade ago at Stanford and the University of Washington, watching the market transition from offshore click-farms to expert precision.
  • Uncurated synthetic data creates a "synthetic slop cannon" when foundation models attempt to train on unverified synthetic outputs without human domain guardrails.
  • True Recursive Self-Improvement (RSI) requires combining specialized AI models with verified human domain experts in continuous feedback loops.
  • Ratner formalizes this industry evolution as the Data 1.0 to Data 2.0 Shift.

The Data 1.0 to Data 2.0 Shift

Alex Ratner spent ten years watching AI teams run the same broken playbook. In the early days of machine learning, models knew so little that throwing massive amounts of cheap, crowdsourced data at them produced steady gains. Today, base foundation models have hit a performance wall where generic data adds zero signal.

To break through modern performance asymptotes, teams must transition across two distinct phases:

  • Data 1.0: Early Learning Curve and Volume Focus

Early on in model development when the baseline model knows little, almost any new bit of information is additive. The rational strategy centers on raw data volume, predominantly solved through staffing, crowdsourcing, and manual human labor factories.

  • Data 2.0: Asymptote Phase and Precision Specialization

As models mature and reach a durable performance asymptote, generalized data loses value. Progress requires high-precision, specialized data targeted at specific model gaps or misalignments, combining deep human domain expertise with specialized AI feedback loops rather than uncurated synthetic generation.

Ratner rejects the Silicon Valley fantasy that models can simply train themselves on their own synthetic output indefinitely. As Ratner notes, “You can't get this hard data with just staffing or sourcing human hours. You can't get it with just trying to have the model that you're trying to sell to that provider generate all the data, like a synthetic slop cannon. You have to actually use more sophisticated blends of human expertise and specialized AI.”

Instead of raw synthetic volume, Snorkel builds an engine for recursive improvement where specialized models speed up human experts, while those experts' corrections continually improve the underlying models.

When This Works (and When It Doesn't)

Ratner's model applies across all learned autonomous systems, including large language models, autonomous vehicles, and physical robotics. It kicks in the moment baseline models achieve basic competence and need expert reasoning to capture the final fractions of accuracy.

This approach is complete overkill for early-stage prototypes. If your product is at the baseline stage where the model cannot parse basic user intent or follow standard syntax, you do not need five-hundred-dollar-an-hour radiologists or corporate lawyers in your loop. In the Data 1.0 phase, off-the-shelf public datasets and basic crowdsourced labels are cheaper and get the job done faster. Save expert feedback loops for when standard training stops moving your core metrics.

What to Do With This

Audit your model evaluation dataset this week. Pull the last 50 production failures where your fine-tuned model hallucinated or gave a bad answer. Group them by failure type.

If the errors stem from basic formatting or general reasoning failures, you are still in Data 1.0; clean your generic prompts. If the errors stem from subtle domain nuances that an entry-level worker would miss, stop buying generic offshore labels. Hire two experienced domain specialists for a three-hour working session. Have your internal model generate candidate answers, let the specialists annotate the precise reasoning errors, and feed those corrected traces back into your specialized training set.