InsightfulDiscussion

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

The a16z Show39m 45s

VALS, an independent AI evaluation company, addresses the gap where public benchmarks fail to accurately measure model capabilities—evidenced by Meta's Llama 4 underperforming on private benchmarks while excelling on public ones. The podcast discusses how third-party evaluators are essential for both labs seeking credible performance proof and enterprises needing ROI justification for AI spending, while also exploring the role of standardized evaluations in policy and geopolitical AI governance.

Summary

The episode features Ben Horowitz from a16z and Ryan Krishnan, founder and CEO of VALS, discussing the critical need for independent AI model evaluation. Krishnan explains that VALS was founded in 2024 after discovering that public benchmarks were insufficient for measuring model progress, especially as multiple capable models entered the market beyond OpenAI. The conversation reveals a fundamental conflict of interest: labs building internal benchmarks to drive their own development cannot credibly report independent performance metrics, analogous to companies auditing themselves—a problem illustrated by Meta's Llama 4, which showed strong public benchmark results but underperformed on VALS' private benchmarks.

The discussion covers VALS' operational challenges, including the ability to run comprehensive evaluations within a six-hour window before model releases, combining both automated infrastructure (their system called "Steve") and human effort. Krishnan explains that evaluations have evolved from simple single-answer tasks to complex agentic workflows requiring infrastructure that can handle hours-long or even multi-week evaluation periods.

Beyond labs, the conversation emphasizes an underappreciated crisis in enterprise: companies are misvaluing AI intelligence across the stack. Krishnan shares an example of a Fortune 10 company allocating $100 per employee daily for AI tool usage, creating perverse outcomes where engineers work most productively during 4-6 PM when rate limits reset. The company later spent $1.5 million in tokens over a month—10x their engineering salary budget—necessitating tools like VALS' ValSmith product to identify cost-optimal model choices for specific tasks and code repositories.

The evaluation framework itself faces challenges mirroring human assessment: just as there's no agreed standard for measuring human intelligence (IQ, EQ, Big Five), AI evaluation requires making implicit professional distinctions explicit. Horowitz draws an analogy to the MPAA's classification system, suggesting evaluation standards will necessarily be fuzzy but will develop through industry consensus over time.

On the policy front, Krishnan argues that government should set and enforce rules while third parties like VALS verify compliance, mirroring how the government sets nuclear policy while inspectors verify adherence. The relationship requires government agencies to specify capabilities they want to prevent (biohacking, cyberattacks) and independent evaluators to test whether models exhibit those capabilities and can be prompted to misuse them.

The transcript concludes with discussion of how benchmarks must continuously evolve—what Krishnan calls "always a higher peak"—retiring saturated benchmarks and creating new ones reflective of current world states, similar to how lawyers and doctors must recertify. Finally, the conversation addresses geopolitical dimensions: as sovereign AI development becomes inevitable, a shared evaluation language will be necessary for international verification of AI capabilities and risks, potentially enabling verification comparable to nuclear arms agreements.

About this episode

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination.

Key Insights

  • Labs cannot credibly self-report model performance because they have incentive to optimize benchmarks toward their own products, similar to the Enron auditing failure where mixed incentives allowed pay-to-pass dynamics
  • Meta's Llama 4 demonstrated the distortion problem: it showed impressive results on public benchmarks where questions and rubrics are open-source, but significantly underperformed on VALS' private, higher-signal benchmarks
  • Enterprise AI spending is becoming a critical cost control problem, with one Fortune 10 company spending $1.5 million on tokens in a single month—10x their engineering salary budget—while employees faced arbitrary $100/day usage limits creating perverse productivity schedules
  • Current AI evaluation practices require making implicit professional distinctions explicit (like what differentiates a law associate from partner), which demands codifying real-world workflows and task repositories into evaluations
  • Government agencies excel at setting and enforcing rules but are poorly suited to verify technical compliance over time, requiring division of labor where policy makers specify prohibited capabilities and third-party evaluators test whether models possess and can be prompted toward those capabilities
  • Benchmarks must continuously deprecate and evolve to stay at the frontier of model capabilities, similar to professional recertification requirements for doctors and lawyers, preventing companies from cherry-picking saturated benchmarks where they perform well
  • VALS can evaluate complex agentic workflows (running for days or weeks) within a six-hour pre-release window through distributed infrastructure and AI-assisted automation systems, enabling rapid independent validation without delaying model launches
  • International AI governance will require a shared evaluation language and verification mechanism comparable to nuclear arms treaties, using recursive self-improvement benchmarks as potential common ground for countries to verify each other's AI development paces

Topics

Third-party AI model evaluation and benchmarkingConflict of interest in self-reported model performanceEnterprise AI ROI and cost optimizationEvaluation infrastructure and agentic workflow testingPolicy and regulatory frameworks for AIGeopolitical AI governance and international verificationBenchmark evolution and deprecationRecursive self-improvement measurement

Transcript

Every time a new trillion dollar industry emerges, there's a need for this independent testing group. When Meta released Lama 4 on our held out private benchmarks, the model was actually underperforming. But on all of the major public benchmarks, it was showing incredible capabilities. What's the limit of what you can achieve? And then within that, how are you going about it? In an ideal world, take a frontier model and have it train the next version of itself. But obviously that's very expensive and slow. And so what we're doing is forming a set of proxies for every part of the process it takes to build the next version of the model. Evaluations as become more complex, have…

Full transcript available for MurmurCast members

Sign Up to Access

More from The a16z Show

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.