DiscussionInsightful

Why Medical AI Needs a Referee | Protege's Engy Ziedan

The a16z Show35m 21s

Engy Ziedan, co-founder of Protégé, argues that medical AI faces a critical evaluation gap: models can ace benchmark exams but fail in real clinical settings, and independent oversight is essential because hundreds of millions use AI for health advice with no one ensuring safety or correctness. Protégé is positioning itself as an impartial referee to continuously evaluate AI models in healthcare through real-world data rather than static benchmarks.

Summary

The podcast features Engy Ziedan discussing why medical AI requires robust, independent evaluation mechanisms. Ziedan explains that Protégé was founded on the insight that frontier models are limited by available training data, and the company pivoted from serving startups to working with large foundation models and enterprise customers seeking data for pre-training, mid-training, and now crucially, for benchmarking and evaluation.

A central problem Ziedan identifies is the gap between benchmark performance and real-world clinical utility. While models can score 92% on medical licensing exams, they perform at only 45% on actual clinical tasks. This disconnect occurs because traditional benchmarks rely on multiple-choice medical questions that models may have memorized rather than truly reasoning through. Additionally, benchmarks don't capture the real-world context of clinical decision-making, physician preferences, and the subtle misalignments that are harder to detect than catastrophic failures.

Ziedan argues that evals are economically critical for AI market equilibrium. Unlike Uber or past technologies where value became obvious through user adoption, AI pricing requires understanding what's actually valuable versus valueless. Without accurate evaluations, healthcare systems cannot make informed purchasing decisions, and misaligned AI systems could cause systemic harm similar to the opioid epidemic—except that with AI's speed and agency, problems escalate faster.

The transcript discusses how every AI vendor claims to be best, but without independent evaluation, no one actually knows which models perform better for specific tasks. Ziedan notes that recent competing papers in Nature and arXiv reached opposite conclusions about whether general frontier models or vertical-specific healthcare AI performs better, demonstrating the contamination and methodology problems plaguing current evaluations.

Ziedan also addresses the subtler problem of misalignment beyond catastrophic failures. He gives the example of AI systems used in insurance prior authorization and medical denials that may be optimizing for revenue minimization rather than patient outcomes, creating situations where AI might check a patient's debt before recommending treatment. This type of misalignment is harder to define and detect than outright failure but represents a profound problem in healthcare.

Protégé's solution involves hosting unified evaluations within specific healthcare subnodes (like oncology pathology or clinical trial recommendations), producing comparative reports where multiple vertical AI builders can be fairly evaluated against each other. The methodology remains transparent, competitor identities are concealed until the end, and they've implemented a strict data membrane: no patient data that entered any model's training can be used in benchmarking, requiring them to source entirely new pathology slides and clinical data.

Ziedan argues Protégé is well-positioned for this referee role because they have comparative advantage in healthcare data and can identify precisely which data would improve a model's performance on revealed deficiencies. Crucially, their incentive structure favors impartiality—any loss of trust destroys the evaluation network immediately. He contrasts this with government oversight, which he argues is neither feasible nor fast enough given AI's rapid evolution, and notes the government itself has asked industry collaborators for help identifying evaluation standards.

About this episode

Daisy Wolf and Eva Steinman are joined by Engy Ziedan, co-founder and Chief Scientific Officer of Protege, to discuss why medical AI has a measurement problem, and why scoring well on a benchmark doesn't necessarily mean a model is ready for the hospital. Engy explains why healthcare AI needs independent evaluations that go beyond static exams and measure how models actually perform in real-world clinical workflows. They explore the risks of subtle bias and misalignment, why the same model can rank differently depending on how it's prompted or tested, and what happens as AI becomes more personalized and changes faster than traditional healthcare quality systems can keep up. The conversation also gets into Protege's role as an independent evaluator, how contaminated training data can undermine benchmarks, and why the future of medical AI may require continuous monitoring rather than occasional testing.

Key Insights

  • Ziedan claims that models scoring 92% on medical licensing exams perform at only 45% on actual clinical tasks, demonstrating that benchmark excellence does not predict real-world utility.
  • He argues that catastrophic AI failure is actually easier to prevent than subtle misalignment, because misalignment is difficult to define, harder to detect, and can persist across systems optimizing for conflicting objectives (e.g., profit maximization versus patient care).
  • Ziedan contends that healthcare AI currently operates under no independent oversight despite hundreds of millions of people using AI for health advice—every vendor claims superiority but without impartial evaluation, no one knows which models actually perform best for specific tasks.
  • He asserts that accurate AI evaluation is economically necessary for market equilibrium and pricing efficiency, unlike past technologies like Uber where value became obvious through user adoption; AI requires understanding what is genuinely valuable versus valueless before deployment.
  • Ziedan argues that AI models may be memorizing benchmark answers rather than performing clinical reasoning, and competing academic papers on the same topic reaching opposite conclusions reveals fundamental contamination problems in current evaluation methodologies.
  • He explains that Protégé can identify exactly which training data would improve a model's performance on revealed deficiencies, giving them a comparative advantage as an evaluator beyond simply being impartial.
  • Ziedan claims that static government oversight modeled on value-based purchasing programs is too slow for AI, which evolves and gains agency faster than healthcare systems can evaluate, creating risks equivalent to historical healthcare crises like the opioid epidemic.
  • He asserts that without credentials, examination, and ongoing evaluation for AI systems in healthcare—standards applied rigorously to imported physicians—there will be growing mistrust of AI generally among patients who prioritize health as core human capital.

Topics

Medical AI evaluation gap and benchmark limitationsReal-world clinical performance versus test performanceMisalignment in AI systems and subtle bias detectionEconomic rationale for AI evaluation and market equilibriumProtégé's role as independent AI arbiter in healthcareData contamination in AI benchmarksAsymmetric information in healthcare and AI adoptionTest-time evaluation versus static benchmarking

Transcript

Hundreds of millions of people ask chat GPT questions about their health. Who, if any, is making sure that the answers that are spit out is safe and correct? Models are going to be inhibited in their usefulness by the training data available for them. There is no one that's looking beyond the iceberg of catastrophic failures and misalignment. What is the importance of evals in this industry? No one ever asks, what is the value of Uber? Show me the eval. But in today's AI market, there is a need for the pricing to be accurate. And without understanding really what is valuable and what's value-less technology, this technology does not have a marginal cost of zero. And so…

Full transcript available for MurmurCast members

Sign Up to Access

More from The a16z Show

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.