Key Takeaways

  • US data labeling companies are actively selling services and valuable training data to Chinese AI labs, prompting a sharp debate among tech investors.
  • David Sacks argues that data labeling is largely a commodity, and restricting its sale could needlessly escalate trade tensions with China without securing a decisive US AI advantage.
  • Jason Calacanis counters that specialized, often “PhD-written” data sets are a “secret sauce” that gives China an unpatriotic shortcut to catching up with American AI labs.
  • Brad Gerstner notes that while this transfer irritates Washington, the US’s current lead in the AI race means it faces less scrutiny now. However, this could swiftly change if China closes the gap.
  • The core tension lies in differentiating between generic, easily replicable data labeling work and proprietary knowledge pipelines that truly embed expert-level intelligence.

The Disagreement

Is American data expertise fueling China's AI ambitions, or are we just exporting a basic commodity? That was the core tension on the All-In Podcast as hosts dissected Forbes' investigation into US data labeling startups selling services to Chinese AI labs. The debate split the room.

Jason Calacanis didn't mince words, arguing these services give China an unfair leg up. “They claim US data labeling startups are selling valuable training data to Chinese labs, which in turn is helping them catch up with the US frontier ones,” Calacanis explained. He stressed that these aren't just generic labels; they're often “created by experts here in America who are given like the queries that have errors in them.” For Calacanis, this specialized, high-quality data is America's “big advantage if we weren't sending it there,” essentially a "secret sauce" helping competitors accelerate their progress.

David Sacks pushed back hard against this view, asserting that most data labeling is not strategically vital. “My sense of data is that it's largely a commodity. I mean, data labeling certainly is,” Sacks stated. He contended that restricting this trade wouldn't stop China, but merely provoke a trade war. “If you basically tell them that they can't use data labeling, I guarantee you there's no shortage of labor in China that they can use to do the data labeling,” Sacks said, suggesting China could easily replicate the effort if needed. He challenged Calacanis to prove any true "proprietary" element was being sold.

Brad Gerstner added a political layer, acknowledging that this practice definitely “will irritate people in Washington.” He noted that current US dominance in AI means less immediate backlash. But he warned that this sentiment, along with other tech exports like chips, could shift dramatically "if China closes the gap" and Washington feels America is making it “too easy on the Chinese labs to catch up.”

Who's Right (and When They're Wrong)

Both Calacanis and Sacks grasp a piece of the truth here, but their arguments apply to different kinds of data labeling. Sacks is absolutely right that basic data labeling — like identifying cats in images or transcribing simple audio — is a commodity. There's nothing proprietary about it, and any country with sufficient labor can build those datasets internally without much trouble. Restricting this type of commodity export would indeed be a futile gesture, mostly serving to sour trade relations.

However, Calacanis is right when he talks about "secret sauce" data. The true strategic value isn't in the raw act of labeling, but in the expertise embedded within the labeling process. Data sets crafted by PhDs, specialized domain experts, or through sophisticated, proprietary pipelines that address complex queries and edge cases, are not commodities. These are knowledge assets that capture nuanced human understanding or specific problem-solving methodologies. Selling this kind of specialized data or the intellectual capital behind its creation offers a real, quantifiable advantage to recipients, shortening their development cycles and improving model performance in critical areas. Sacks himself conceded, “if there's something truly proprietary here, I don't want us to sell our secret sauce to China.”

The distinction is critical: generic data labeling is a replaceable cost; expert-driven data creation is a strategic asset.

What to Do With This

If you're a founder building an AI product or running a data service, audit your current data strategy with a sharp eye. Don't confuse cheap labor for a strategic moat. Ask yourself: Is the data you're buying, selling, or generating a simple commodity that any competitor can replicate with enough effort? Or does it embed truly unique human expertise, proprietary methodologies, or nuanced intellectual property that would be difficult for an adversary to rebuild from scratch? If your "data advantage" relies on easily replicable inputs, you don't have a defensible advantage. Focus immediately on creating data pipelines or expert-driven labeling processes that yield truly proprietary assets, not just commodity inputs. These are the assets Washington will care about, and more importantly, they are the ones that will actually protect your business long-term.