Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
VALS, an independent AI evaluation company, addresses the gap where public benchmarks fail to accurately measure model capabilities—evidenced by Meta's Llama 4 underperforming on private benchmarks while excelling on public ones. The podcast discusses how third-party evaluators are essential for both labs seeking credible performance proof and enterprises needing ROI justification for AI spending, while also exploring the role of standardized evaluations in policy and geopolitical AI governance.
Summary
The episode features Ben Horowitz from a16z and Ryan Krishnan, founder and CEO of VALS, discussing the critical need for independent AI model evaluation. Krishnan explains that VALS was founded in 2024 after discovering that public benchmarks were insufficient for measuring model progress, especially as multiple capable models entered the market beyond OpenAI. The conversation reveals a fundamental conflict of interest: labs building internal benchmarks to drive their own development cannot credibly report independent performance metrics, analogous to companies auditing themselves—a problem illustrated by Meta's Llama 4, which showed strong public benchmark results but underperformed on VALS' private benchmarks.
The discussion covers VALS' operational challenges, including the ability to run comprehensive evaluations within a six-hour window before model releases, combining both automated infrastructure (their system called "Steve") and human effort. Krishnan explains that evaluations have evolved from simple single-answer tasks to complex agentic workflows requiring infrastructure that can handle hours-long or even multi-week evaluation periods.
Beyond labs, the conversation emphasizes an underappreciated crisis in enterprise: companies are misvaluing AI intelligence across the stack. Krishnan shares an example of a Fortune 10 company allocating $100 per employee daily for AI tool usage, creating perverse outcomes where engineers work most productively during 4-6 PM when rate limits reset. The company later spent $1.5 million in tokens over a month—10x their engineering salary budget—necessitating tools like VALS' ValSmith product to identify cost-optimal model choices for specific tasks and code repositories.
The evaluation framework itself faces challenges mirroring human assessment: just as there's no agreed standard for measuring human intelligence (IQ, EQ, Big Five), AI evaluation requires making implicit professional distinctions explicit. Horowitz draws an analogy to the MPAA's classification system, suggesting evaluation standards will necessarily be fuzzy but will develop through industry consensus over time.
On the policy front, Krishnan argues that government should set and enforce rules while third parties like VALS verify compliance, mirroring how the government sets nuclear policy while inspectors verify adherence. The relationship requires government agencies to specify capabilities they want to prevent (biohacking, cyberattacks) and independent evaluators to test whether models exhibit those capabilities and can be prompted to misuse them.
The transcript concludes with discussion of how benchmarks must continuously evolve—what Krishnan calls "always a higher peak"—retiring saturated benchmarks and creating new ones reflective of current world states, similar to how lawyers and doctors must recertify. Finally, the conversation addresses geopolitical dimensions: as sovereign AI development becomes inevitable, a shared evaluation language will be necessary for international verification of AI capabilities and risks, potentially enabling verification comparable to nuclear arms agreements.
About this episode
a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination.
Key Insights
- Labs cannot credibly self-report model performance because they have incentive to optimize benchmarks toward their own products, similar to the Enron auditing failure where mixed incentives allowed pay-to-pass dynamics
- Meta's Llama 4 demonstrated the distortion problem: it showed impressive results on public benchmarks where questions and rubrics are open-source, but significantly underperformed on VALS' private, higher-signal benchmarks
- Enterprise AI spending is becoming a critical cost control problem, with one Fortune 10 company spending $1.5 million on tokens in a single month—10x their engineering salary budget—while employees faced arbitrary $100/day usage limits creating perverse productivity schedules
- Current AI evaluation practices require making implicit professional distinctions explicit (like what differentiates a law associate from partner), which demands codifying real-world workflows and task repositories into evaluations
- Government agencies excel at setting and enforcing rules but are poorly suited to verify technical compliance over time, requiring division of labor where policy makers specify prohibited capabilities and third-party evaluators test whether models possess and can be prompted toward those capabilities
- Benchmarks must continuously deprecate and evolve to stay at the frontier of model capabilities, similar to professional recertification requirements for doctors and lawyers, preventing companies from cherry-picking saturated benchmarks where they perform well
- VALS can evaluate complex agentic workflows (running for days or weeks) within a six-hour pre-release window through distributed infrastructure and AI-assisted automation systems, enabling rapid independent validation without delaying model launches
- International AI governance will require a shared evaluation language and verification mechanism comparable to nuclear arms treaties, using recursive self-improvement benchmarks as potential common ground for countries to verify each other's AI development paces
Topics
Transcript
Every time a new trillion dollar industry emerges, there's a need for this independent testing group. When Meta released Lama 4 on our held out private benchmarks, the model was actually underperforming. But on all of the major public benchmarks, it was showing incredible capabilities. What's the limit of what you can achieve? And then within that, how are you going about it? In an ideal world, take a frontier model and have it train the next version of itself. But obviously that's very expensive and slow. And so what we're doing is forming a set of proxies for every part of the process it takes to build the next version of the model. Evaluations as become more complex, have…
Full transcript available for MurmurCast members
Sign Up to AccessMore from The a16z Show
OpenAI Researchers on the Future of Mathematical Reasoning
OpenAI researchers discuss how AI models are making progress on long-standing mathematical problems by combining literature knowledge, executing complex proofs with precision, and exploring multiple approaches without human cognitive biases. They present case studies in sphere packing, coding theory, and group theory, arguing that AI's ability to persist through difficult problems and leverage symmetry properties is fundamentally changing what mathematics gets solved and how it's practiced.
Can Open Source Keep AI Power From Concentrating?
Lucas Kaiser, co-author of the Transformer paper, discusses how AI power is currently concentrating in large companies due to the resource-intensive nature of current technology, but argues this is not inevitable. He believes research breakthroughs in algorithms and training methods could enable smaller players and distributed models to compete effectively.
Your AI Doctor Is Coming | Julie Yoo
Julie Yoo, a healthcare investor at Andreessen Horowitz, argues that AI will benefit healthcare more than any other industry because healthcare has historically underinvested in technology, allowing it to leapfrog legacy systems and adopt AI-native solutions directly. She identifies major opportunities in consumer healthcare, AI-native care delivery, robotics, and new payment models, predicting a future where individuals have personalized AI doctors available continuously.
Aaron Levie on Why Open AI Wins
Aaron Levie, CEO of Box, discusses why open-weight AI models benefit the entire AI ecosystem rather than threatening frontier labs, argues that the economic value in AI accrues to inference infrastructure rather than model weights, and explains why model routing will become the default enterprise AI strategy.
Fei Fei Li: The Race to Build World Models For AI
World Labs launched Atlas, a new world model built on next-view prediction that unifies 3D reconstruction and generation. The model can create spatially-grounded video frames from sparse input images (as few as three), reducing the data requirements for 3D scene capture by 50-100x, with applications ranging from creative content to robotics simulation.