Key Takeaways
- Static persona prompts fail because frontier models collapse into stereotypes rather than reflecting actual human choices.
- Simile AI achieves 85% fidelity by building digital twins from deep interviews, transaction records, and randomized trials across tens of thousands of consenting participants.
- Wealthfront used Simile to run multimodal tests directly on Figma mockups, studying how digital twins react to visual designs before writing frontend code.
- Enterprise customers like Gallup, CVS, and Deloitte use domain-agnostic models to test product concepts and forecast earnings call reactions.
The Trap of Prompted Personas
Most teams testing product ideas with AI make the same mistake. They write a prompt telling ChatGPT to act like a 35-year-old accountant with two kids. The model spits back generic platitudes because prompt engineering only pulls from statistical averages across web text. It reflects how people write on Reddit, not how individuals spend money or react to software.
Joon Sung Park saw this limit while leading Stanford's Generative Agents project. At Simile AI, he replaced prompt-based personas with continuous data pipelines. His team recruits tens of thousands of consenting people, gathering deep interviews, observational transaction records, and randomized experiments. The result is a population of grounded digital twins that reproduce human actions with 85% fidelity.
Testing Figma Mockups Before Shipping Code
Traditional market research relies on surveys or focus groups that take weeks to schedule. Even worse, participants often say what they think researchers want to hear. Simile shifts product validation from post-build surveys to pre-build simulations.
Wealthfront became an early test case when they pushed Simile beyond text questions. “Wealthfront was one of the first customers that wanted to actually do product testing that goes beyond just asking people what they think,” Park explained. “There, really what we had to do was reason about multimodal input.”
Instead of asking simulated users abstract questions about financial habits, Wealthfront ran digital twins directly against Figma mockups. The models analyzed interface layouts, visual hierarchy, and copy choices, simulating how different demographic cohorts react to specific screens. If an onboarding step creates friction or a pricing tier confuses users, the simulation flags it before engineers write a single line of code.
Replacing the $100B Research Panel
Market research is a massive industry, but Park views traditional consumer panels as too narrow. Legacy panels sell static answers to pre-written questions. When enterprise clients like Deloitte, CVS, or Gallup want to forecast market shifts or earnings calls, static surveys fail because they cannot adapt to fluid conditions.
“Market research is a hundred billion dollar industry,” Park noted. “But simulation is not a tool for market research. Simulation is a tool for human decision-making.”
When you capture behavioral dynamics instead of survey answers, the same digital twin can evaluate a retail loyalty program, predict churn on a fintech app, or simulate public reactions to a policy change. The goal is to bring missing voices into executive discussions before high-stakes bets go live.
What to Do With This
Audit your product testing pipeline this week. Stop running synthetic tests with generic prompts like "Act like a target customer." Instead, map three real customer personas to specific behavioral data: their last three bank transactions, exact onboarding drop-off points, and direct support tickets. Feed those concrete histories into your test prompts to eliminate generic model responses.