Key Takeaways

  • Thinking Machine Lab's new 'Inkling' model, billed as an open-weights competitor to OpenAI and Anthropic, landed right in the middle of a heated debate about AI model originality.
  • John Coogan claimed Inkling was unique as the "only openweight model" not distilled from OpenAI or Anthropic, distinguishing it from popular models like Kimmy and Deepseek.
  • This claim immediately hit a snag when Tyler Cosgrove highlighted Inkling's own blog post, which stated their model used “synthetic data generated by open weight models including Kimmy K2.5”—models Coogan previously identified as distilled.
  • Anthropic is actively fighting back against this practice, reportedly shutting down “millions of [distillation] accounts per week,” signaling a severe, ongoing challenge to IP in AI.
  • The entire discussion casts a shadow on the competitive landscape, suggesting that many "open-source" AI models might rely on a complex, often uncredited, chain of training data derived from closed-source leaders.

The Disagreement

The rollout of Thinking Machine Lab's Inkling model was supposed to be a win for open-source AI. John Coogan, hosting TBPN, kicked things off by positioning Inkling as a direct challenge to the closed-source giants. “Thinking machines lab first, uh, the first model is an open weights model designed to chip away at the lead of OpenAI and anthropics, says the Wall Street Journal,” Coogan noted. He then leaned into a bold claim: Inkling was somehow purer, more original than its open-source peers. “This is the only openweight model that that's trained without distilling for OpenAI from OpenAI or anthropic. Kimmy distills, GLM distills, Quen distills, Neotron distills, Kimmy and Deepseek, which count,” Coogan stated, directly naming other models he believed were derived.

But that purity claim didn't last long. Guest host Tyler Cosgrove quickly pulled the thread, reading directly from Inkling's own blog post. “in the blog post they say to bootstrap post training we ran an initial supervised fine-tuning on synthetic data generated by open weight models including Kimmy K2.5,” Cosgrove revealed. The tension was immediate. If Inkling used synthetic data from Kimmy—a model Coogan himself just labeled as distilled—how "undistilled" could Inkling truly be? This exchange highlighted a critical, often obscured, aspect of AI development: the hidden dependencies and complex lineage of training data, especially in the "open-source" space.

Who's Right (and When They're Wrong)

Both Coogan and Cosgrove revealed parts of a complex truth. Coogan's initial enthusiasm for Inkling's stated non-distilled status highlights the desire for truly original open-source AI. It speaks to the ideal of independent innovation chipping away at the lead of large tech firms. The industry wants models that are genuinely built from the ground up, avoiding the sticky legal and ethical questions of "copying."

Cosgrove, however, was right to challenge the claim based on Inkling's own disclosures. His point illustrates how quickly the definition of "original" blurs in AI. When models like Kimmy are identified as distilled, then Inkling uses data from Kimmy, the chain of originality gets murky. This isn't necessarily about malice, but about the practical realities of training advanced AI models. Developers often bootstrap by using synthetic data, and if that synthetic data comes from models that themselves are derived, the "undistilled" label becomes a technicality. The debate reveals a silent battle being waged: Anthropic, for instance, is fighting hard against this very practice, with Coogan reporting they are "shutting down distillation accounts on the order of millions accounts of per week. That is crazy scale." This suggests that even if a model intends to be original, the ecosystem makes it incredibly difficult.

The implication is clear: even seemingly independent open-source models may carry a hidden "genetic code" from larger, closed-source incumbents. This works for rapid iteration and making powerful AI accessible, but it fails when originality, legal defensibility, or competitive advantage are paramount. If your business depends on a truly unique model, you need a clearer line of sight into your entire data pipeline, not just the model you're fine-tuning.

What to Do With This

If you're building with or on open-source AI models, don't just take "not distilled" at face value. This week, audit your model's data lineage: trace the synthetic data, the fine-tuning datasets, and every foundational model back to its source. If you find dependencies on models accused of distillation, develop a mitigation plan for potential IP risks, or shift your strategy to training on truly proprietary data to secure your competitive edge.