Key Takeaways

  • Jun Song Park, CEO of Simily, argues modern AI companies need defensible data strategies that go beyond mere observation.
  • Clients don't just want AI to predict the future; they want models that can reason about causal mechanisms to actively shape it.
  • Simily builds its intelligence by collecting unique data from randomized control trials (RCTs) and A/B testing with everyday people.
  • This method focuses on understanding what people do in response to specific interventions, rather than just what they say or passively exhibit online.

The Causal Data Method: Why Simily Rejects Mere Prediction

Forget your usual data lake. Jun Song Park, co-founder and CEO of Simily, lays down a sharp truth for ambitious builders: “for AI companies of this generation, you need to have an interesting data strategy that's going to be defensible.” Most AI today chews through vast amounts of observational data, predicting what might happen. But Park points out a critical flaw in this approach: “No one really cares about prediction... What people actually care about is they want to shape the future.”

Think about it: if you're Starbucks and your AI predicts Frappuccino sales will tank next quarter, that's just a warning. It doesn't tell you how to stop it. As Park puts it, “doesn't really help them to know that your Frappuccino sales is going to tank in two quarters... What they want to know is how can we prevent it? What do we need to do now to change the future?” To answer how to change the future, AI needs to understand cause and effect, not just correlation.

Beyond Web Data: How Simily Builds Its Causal Engine

Simily's data strategy is a direct attack on this problem. Instead of scraping web data – which largely captures what people say or passively do – they engineer scenarios to see what people actually do under specific conditions. “The kind of data that we care deeply about is a lot of randomized control trials. We actually run a lot of AB testing,” Park explains. This means actively intervening and observing the outcomes, rather than just passively watching.

Who are these tests run on? Not expert programmers or scientists. “We go after people like us, like everyday people and living their everyday life,” says Park. By simulating scenarios and varying conditions for a broad base of ordinary users, Simily collects rich behavioral and transactional data. This unique dataset allows their models to learn not just what happened, but why it happened, and, crucially, what would happen if a different action were taken. “We show the models, imagine people have done this versus that. This is how their behaviors will actually change. That becomes a core part of our training asset.”

Where This Breaks Down

While powerful, Simily's approach isn't a silver bullet for every founder. Running randomized control trials (RCTs) and A/B tests to collect causal data is inherently more resource-intensive than simply acquiring pre-existing observational datasets. It demands careful experimental design, participant recruitment, and a longer data collection cycle. This can be slower and more expensive, especially for early-stage startups with tight budgets or for AI applications that don't directly involve human behavior. Moreover, the generalizability of small-scale RCTs with "everyday people" might not always scale to the vast, diverse populations required for some enterprise solutions, potentially limiting its scope where mass market patterns are paramount.

What to Do With This

Stop optimizing for more data; start optimizing for better data. This week, identify one critical customer behavior you're trying to influence – not just predict. Instead of A/B testing a UI tweak for conversion rate, design a mini-RCT to understand why users drop off at a specific point. Offer 50 users one incentive to complete a step, and 50 users a different one. Track the causal impact, not just the correlation. This shifts your thinking from forecasting trends to engineering outcomes.