Key Takeaways

  • Micro1 offered Sam Parr's executive network Hampton $800,000 to plug directly into its private Notion, Slack, Gmail, and Google Drive workspaces.
  • Frontier AI labs have exhausted the crawlable public web, shifting the battle for model performance toward private corporate reasoning and human workflows.
  • AI companies like Anthropic have bought physical out-of-print books and built custom scanners to feed rare knowledge into model training runs.
  • Micro1 acts as a data broker for the roughly eight major LLM developers, turning proprietary company logs into training pipelines.

The $800,000 Offer for Hampton's Internal Workspace

Sam Parr recently saw the AI data rush land directly in his inbox. Micro1, a fast-growing data startup, approached his executive community company, Hampton, with a clear financial pitch.

“Micro One is a company,” Parr explained. “Basically what they've been doing is they've been going to privately held companies like my company, Hampton, and they say, 'Hey, if you plug in your notion, your Slack, your Gmail, your Drive, and a bunch of other information, and they will give you $800,000.'”

The offer highlights how desperate AI labs have become for high-quality human interactions. Common Crawl, Reddit threads, and Wikipedia articles have already been ingested across multiple generations of model training. The remaining frontier of human reasoning lives behind private enterprise logins: executive decision-making threads, messy operational Slack debates, customer onboarding documents, and internal strategy wikis.

Physical Scanners and the Scramble for Private Data

Shaan Puri pointed out that model builders cannot keep recycling identical training datasets if they want their next release to beat the competition.

“All the models so far have been trained on the open internet,” Puri said. “So what was publicly available, and they don't want to use the same input for each model train that they do.”

To escape the public web bottleneck, labs are going to extreme lengths. Puri noted that Anthropic began acquiring out-of-print books: “They built all these scanners that will scan page after page after page of these books.” When physical print libraries are not enough, data brokers step in to aggregate internal business operations.

“Their business model is that they want all types of data that aren't available to the public,” Parr said. “And they sell it to the large... when people say large LLMs, I imagine there's only like eight of them.”

The financial upside for the middlemen is massive. Parr observed that Micro1's founder is around 25 years old and already a paper billionaire brokering these proprietary data deals. “It's just crazy how much wealth is being created,” Parr said, “and guys like you and I, we're just getting the crumbs.”

What to Do With This

Audit your company's proprietary data footprint this week. Check your customer agreements to see if internal communications or user-generated workflows are restricted from AI training pipelines before you plug third-party aggregators into your Slack or Notion workspaces. If you sit on unique operational datasets, price that data as an asset rather than giving it away through unvetted software integrations.