InsightfulDiscussion

20VC: The Best AI Companies Have Unique Data Acquisition Strategies | Will Simile Kill Kalshi, Polymarkets and NASDAQ | How to Sign Fortune 500 Companies As Customers in Weeks with Joon Sung Park, Simile

Joon Sung Park, founder of Simile, discusses how his company is building a foundation model of human behavior to create simulations that predict how people and markets will respond to future scenarios. The company has raised $300 million in six months and is gaining rapid enterprise adoption by helping Fortune 500 companies avoid costly mistakes and test decisions before implementation.

Summary

Joon Sung Park founded Simile, a simulation company that models human behavior using AI agents with memory, planning, and reflection capabilities. The company's origins trace to a 2023 Valentine's Day project where Park created a virtual town with 25 NPCs that self-organized and planned parties, demonstrating that large language models could extract realistic human behaviors from training data. This early work established the technical foundations for memory systems (using markdown text files) and reflection mechanisms that allow agents to synthesize experiences and develop personalities.

Simile positions itself as building a "foundation model of human behavior" rather than a simulation company—this foundation can be used to create individual simulations, subpopulation simulations, and eventually entire ecosystem simulations. The key differentiator from general-purpose LLMs like those from OpenAI is that Simile optimizes for behavioral realism and human-like mistakes rather than perfect rationality. Park emphasizes the gap between what people say and what they do (say-to-give problem), arguing that most LLM training data reflects stated behavior, not actual behavior.

The company's data strategy centers on collecting behavioral data—transaction data, observational data, and crucially, randomized control trial results. Park argues that observational data is good for prediction but inadequate for the counterfactual reasoning his customers need. Enterprises don't want to know what will happen; they want to know how to prevent negative outcomes or encourage positive ones. This requires causal mechanisms and counterfactual understanding, not just correlation.

On the enterprise side, Simile has achieved remarkable sales velocity, closing deals with Fortune 500 companies like CVS within three months rather than the traditional year-plus enterprise sales cycles. Park attributes this to acute customer pain around slow experimentation, limited budgets, and inability to test ideas. Early validation came when Simile ran simulations that replicated three-to-six-month consulting studies in two minutes. The company has demonstrated 85% prediction accuracy compared to human replication rates and is building what the industry calls "synthetic panels"—AI-powered market research replacing human panels.

On team building, Park emphasizes hiring for "common denominator" success—people who were the reason projects succeeded and can reinvent themselves across domains. He also values finding leaders with seemingly contradictory superpowers, such as CMOs who are both deeply analytical and creative. His co-founder Lainey exemplifies the ideal of being simultaneously paranoid short-term and religiously long-term optimistic—a balance that drives execution without complacency.

Regarding funding, Simile raised $100 million six months ago, then raised an additional $200 million led by Grelock Partners and insiders at Index Ventures. Park noted he didn't initially seek the second round but accepted it to increase compute spend and data collection, allowing meaningful acceleration of research progress. He also reflected on unexpected lessons from fundraising: the genuine mentorship value VCs can provide and how market interest moves faster than expected.

On data collection challenges, Park states that data acquisition is the hardest element of building simulation models. The company sources everyday people (not experts) who are representative of real-world populations and conducts extensive interviews asking about life stories, major decisions, and formative experiences. The question of how much data is needed depends on segmentation; roughly 1,000 people provides statistical significance for narrow populations, but Simile aims to represent entire populations to support on-the-fly filtering queries.

The data flywheel mechanism works through continuous validation against ground truth. Rather than waiting years for outcomes, Simile generates tens of thousands of daily hypotheses and validates them as the world unfolds. This provides learning signals superior to traditional AI training in coding (which has clear accept/reject signals) because the entire world becomes the ground truth.

Park envisions simulation reaching a stage where single simulation sessions cost $10-100 million to run but generate such valuable insights that enterprises will pay accordingly. He sees simulation as eventually addressing "wicked problems" requiring collective action across stakeholders with different incentives, such as climate change or democratic governance.

On the competitive landscape, Park addresses whether Simile threatens prediction markets like Kalshi and Polymarket by noting the difference: prediction markets forecast what will happen, while Simile reveals how and why it happens and how to change it. He's also careful about ethical considerations, viewing simulation as a powerful technology with potential for misuse and emphasizing that its best application is ensuring diverse perspectives are heard in decision-making.

Park describes his background as an atypical researcher—he had no research experience during undergrad and lived in a garage after college. He credits Professor Mary Wooldridge at Stanford with taking an undeserved chance on him by spending a full morning providing guidance and introductions that launched his research career, a kindness he remains deeply grateful for.

About this episode

<p dir="ltr">Joon Sung Park is the Founder and CEO of Simile, the AI simulation company building foundation models of human behaviour; allowing companies to test how real people may think, decide and act before making a decision in the real world. Simile has now raised $300 million in total, including a $200 million Series B announced last week at a $2 billion valuation, led by Greenoaks and Index Ventures.</p> <p dir="ltr">AGENDA:</p> <p dir="ltr">00:00 We Will Pay $100M for a Single Query on Some Models</p> <p dir="ltr">10:00 Why Stock Markets May Not Exist in 5 Years Time</p> <p dir="ltr">15:00 The Best Companies All Have Unique Data Acquisition Strategies</p> <p dir="ltr">19:00 The Best AI Companies Have Clear and Fast Reward Functions</p> <p dir="ltr">24:00 How We Sign Fortune 500 Companies for $10M Contracts in Weeks</p> <p dir="ltr">32:00 Does Similie Kill Kalshi and Polymarket? Prediction vs Changing the Future</p> <p dir="ltr">42:00 Inside Similie's $300M Raise; What Every Founder Needs to Know</p> <p dir="ltr"> </p> <p> </p>

Key Insights

  • Large language models trained on web data capture what people say, not what they do, creating a fundamental gap that Simile addresses by collecting actual behavioral data including transaction records and randomized control trial results.
  • Simile's core customers don't primarily care about prediction; they care about prevention and counterfactual reasoning—knowing how to change outcomes rather than simply forecasting them, which requires causal mechanisms rather than correlational observations.
  • Enterprise sales cycles for Simile's simulation technology moved 3-5x faster than traditional enterprise software, with Fortune 500 companies closing deals in three months because the pain point (slow, expensive experimentation) was more acute than anticipated.
  • The company built a data flywheel where tens of thousands of daily hypotheses are continuously validated against ground truth events, allowing faster learning than traditional AI training approaches that rely on explicit accept/reject signals.
  • Simile's memory system for AI agents uses simple markdown text files parsed by language models, with reflection mechanisms that periodically synthesize memories into higher-level patterns, enabling agents to develop personalities and long-term understanding.
  • Park views researchers as motivated primarily by vision and societal impact rather than salary alone, noting that observing companies like OpenAI and Anthropic grow rapidly from laughingstock to trillion-dollar valuations creates confidence in ambitious technical visions.
  • The company deliberately targets representativeness of everyday people in its data collection rather than experts, asking about life stories and formative decisions to capture the causal foundations of human behavior.
  • Park argues that individual prediction accuracy matters less than understanding emergent ecosystem-level effects when multiple people interact, making multi-agent simulation a fundamentally different capability from single-point predictions.
  • Synthetic panels created by Simile are positioned to replace traditional human market research panels entirely, with the market expected to tip within three years due to cost, speed, and scalability advantages.
  • The company raised $200 million in a secondary round not out of necessity but to increase compute spend and data collection, reflecting the strategic insight that more input (data and compute) can meaningfully accelerate research progress.
  • Park identifies the key differentiator between Simile and general-purpose LLMs as Simile's focus on representing human subjectivity, values, preferences, and biases rather than building perfectly rational superintelligent machines.
  • The vision for simulation extends beyond commercial use to addressing wicked problems like climate change and democratic governance by simulating multi-stakeholder decision-making processes with conflicting incentives.

Topics

Human behavior simulation using AI agentsData collection strategy for behavioral modelsEnterprise sales and product-market fitCounterfactual reasoning vs. predictive analyticsFoundation models for human behaviorTeam building and founder mindsetsVenture funding dynamics and accelerationSay-to-give gap in human behaviorSynthetic panels replacing market researchLong-term vision for simulation technologyEthical considerations in AI simulationCompetitive positioning against prediction markets

Transcript

My fundamental thesis here is for AI companies of this generation, you need to have an interesting data strategy that's going to be defensible. I think there's a world in which in about two, three years, we're running a single simulation session that people will pay $100 million for it. People live through different stages in their life, and they have different careers, different jobs. At each stage of their life, were they the reason why that thing was successful? If you squint, were they the common denominator? This is 20 VC with me, Harry Stebbings, and I am too old and bored of standard podcast intros. So let me tell you a story. Shardul Shah, one of the best…

Full transcript available for MurmurCast members

Sign Up to Access

More from The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.