InsightfulDiscussion

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

Diane Penn, Anthropic's first technical PM, discusses how the company evolved from an underdog to a $50B+ ARR leader by combining frontier models with exceptional products. She explains how PMs now use 'evals as PRDs,' prioritize experimentation over perfect planning, and must stay deeply hands-on with AI technology to guide teams effectively.

Summary

Diane Penn joined Anthropic in 2023 as its first technical product manager when the company had only five engineers and no released model. Despite initial skepticism that OpenAI had already won the race, Anthropic found its identity through strategic focus on coding as a differentiating use case. Penn traces key inflection points: Opus 3's launch (early 2024) proved Anthropic could build frontier models; the Golden Gate Bridge feature showcased how research could be brought to users quickly; and Opus 4.5 (2024) demonstrated the power of pairing advanced models with exceptional products like Cloud Code.

Penn describes a fundamental shift in product management methodology at Anthropic. Rather than traditional PRDs, the team now uses 'evals as the new PRDs'—they identify specific user pain points through transcripts and feedback, write comprehensive test cases that capture both failure scenarios and success criteria, and track model improvements against these evals. This emerged from concrete problems: early users reported Claude couldn't follow schemas properly, so the team created 40 examples of JSON formatting failures, formalized them as an eval, and measured progress across model versions until the problem was solved to near-perfection.

The company operates in what Penn calls 'the exponential,' where each model generation brings discontinuous capability jumps. Rather than predict exactly what's possible, the organization maintains adaptability through continuous evals and experimentation. Labs functions as an internal innovation engine that identifies 'discontinuous large bets' outside the core roadmap—including Cloud Code, Skills, Cloud Design, and MCP—using small, self-driven teams with strong opinions on themes but flexible approaches to prototypes.

Penn emphasizes that product management has fundamentally changed in the AI era. Managers must spend significant time hands-on with the technology, 'sweating tokens as much as pixels.' She uses Claude extensively herself—building custom Skills to improve her management capabilities, consulting it on pricing decisions, and using it to prepare for crucial conversations. She maintains her own point of view by thinking first, then using Claude as a sparring partner that pushes back rather than simply agrees.

The company's culture prioritizes collaboration and experimentation over individual work. Penn describes discovering use cases in shared Slack channels where employees test new versions and iterate rapidly on ideas together. She attributes Anthropic's ability to ship multiple model series per quarter to radical ownership, team collaboration, and deliberate hiring for low-ego, mission-aligned individuals who support each other. She explicitly rejects the notion that PMs are unnecessary when models are powerful; instead, the relentless work of understanding user needs, providing actionable feedback to researchers, and discovering what's possible with new capabilities makes PMs more essential than ever.

About this episode

<p><strong>Dianne Penn</strong> is Head of Product for Anthropic’s AI Research and Labs teams. She joined in 2023 as Anthropic’s first technical product manager, when the entire product team was five engineers, and has since helped ship every model from Claude 2 through Fable, and helped incubate Claude Code, MCP, Skills, computer use, tool use, and reasoning. Before Anthropic, she helped build Alexa’s AI at Amazon and, before that, traded high-yield bonds at JP Morgan Chase.</p><p></p><p><strong>In our in-depth conversation, we discuss:</strong></p><p>1. What Anthropic’s early days were like</p><p>2. The inflection points that turned Anthropic from an underdog into the fastest-growing company in history</p><p>3. How exactly Claude got so good at coding</p><p>4. The eval-driven development loop her team is pioneering</p><p>5. How to find joy in AI when everything is moving this fast</p><p>6. Why Claude’s willingness to push back is key to its success</p><p>7. Where human judgment remains irreplaceable</p><p>—</p><p><strong>Brought to you by:</strong></p><p><a href="https://workos.com/lenny" target="_blank"><strong>WorkOS</strong></a>—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more</p><p><a href="https://mercury.com/command?utm_source=lennys&#38;utm_medium=sponsored_newsletter&#38;utm_campaign=26q3_brand_campaign" target="_blank"><strong>Mercury</strong></a>—Radically different banking, now with Command</p><p>—</p><p><strong>Episode transcript: </strong><a href="https://www.lennysnewsletter.com/p/anthropics-first-technical-pm-on" target="_blank">https://www.lennysnewsletter.com/p/anthropics-first-technical-pm-on</a></p><p>—</p><p><strong>Archive of all Lenny's Podcast transcripts: </strong><a href="https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&#38;st=ahz0fj11&#38;dl=0" target="_blank">https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&amp;st=ahz0fj11&amp;dl=0</a></p><p>—</p><p><strong>Where to find Dianne Penn:</strong></p><p>• LinkedIn: <a href="http://linkedin.com/in/dianne-na-penn" target="_blank">linkedin.com/in/dianne-na-penn</a></p><p>—</p><p><strong>Where to find Lenny:</strong></p><p>• Newsletter: <a href="https://www.lennysnewsletter.com" target="_blank">https://www.lennysnewsletter.com</a></p><p>• X: <a href="https://twitter.com/lennysan" target="_blank">https://twitter.com/lennysan</a></p><p>• LinkedIn: <a href="https://www.linkedin.com/in/lennyrachitsky/" target="_blank">https://www.linkedin.com/in/lennyrachitsky/</a></p><p>—</p><p><strong>In this episode, we cover:</strong></p><p>(00:00) Introduction</p><p>(02:31) Early Anthropic days</p><p>(08:55) Big milestones</p><p>(13:50) Inside the exponential</p><p>(20:02) Token maxing</p><p>(23:30) Anthropic Labs and the incubation model</p><p>(27:30) How the research role works</p><p>(31:35) How to become a top researcher</p><p>(35:18) Frontier model safeguards</p><p>(39:38) Hiring in the AI era</p><p>(44:16) Building an eval set</p><p>(47:48) Evals vs PRDs</p><p>(49:55) The importance of hands-on leadership</p><p>(52:46) Finding joy in AI</p><p>(58:10) How Dianne uses Claude</p><p>(01:01:05) Avoiding overreliance on AI</p><p>(01:03:50) The constitution that makes Claude better</p><p>(01:07:11) AI writing and verification</p><p>(01:11:40) Where human brains will continue to be valuable</p><p>(01:14:10) Navigating AI with kids</p><p>(01:16:26) Alignment, the future of the PM role, and burnout</p><p>(01:21:54) Lightning round and final thoughts</p><p>—</p><p><strong>Referenced:</strong></p><p>• Anthropic: <a href="https://www.anthropic.com" target="_blank">https://www.anthropic.com</a></p><p>• Golden Gate Claude: <a href="https://www.anthropic.com/news/golden-gate-claude" target="_blank">https://www.anthropic.com/news/golden-gate-claude</a></p><p>• Dario Amodei’s website: <a href="https://darioamodei.com" target="_blank">https://darioamodei.com</a></p><p>• Scaling Laws and Interpretability of Learning from Repeated Data: <a href="https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data" target="_blank">https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data</a></p><p>• Tokenmaxxing: How Top Builders Use AI To Do The Work Of 400 Engineers: <a href="https://www.ycombinator.com/library/Pa-tokenmaxxing-how-top-builders-use-ai-to-do-the-work-of-400-engineers" target="_blank">https://www.ycombinator.com/library/Pa-tokenmaxxing-how-top-builders-use-ai-to-do-the-work-of-400-engineers</a></p><p>• Garry Tan on X: <a href="https://x.com/garrytan" target="_blank">https://x.com/garrytan</a></p><p>• Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night | Ben Mann: <a href="https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann" target="_blank">https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann</a></p><p>• Anthropic’s CPO on what comes next | Mike Krieger (co-founder of Instagram): <a href="https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next" target="_blank">https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next</a></p><p>• Introducing Labs: <a href="https://www.anthropic.com/news/introducing-anthropic-labs" target="_blank">https://www.anthropic.com/news/introducing-anthropic-labs</a></p><p>• Louis CK | about airplane Wi Fi: <a href="https://www.youtube.com/watch?v=me4BZBsHwZs" target="_blank">https://www.youtube.com/watch?v=me4BZBsHwZs</a></p><p>• What happens after coding is solved? | Fiona Fung (Manager of the Claude Code and Cowork Teams): <a href="https://www.lennysnewsletter.com/p/building-the-most-ai-pilled-engineering" target="_blank">https://www.lennysnewsletter.com/p/building-the-most-ai-pilled-engineering</a></p><p>• The Anthropic Hive Mind: <a href="https://steve-yegge.medium.com/the-anthropic-hive-mind-d01f768f3d7b" target="_blank">https://steve-yegge.medium.com/the-anthropic-hive-mind-d01f768f3d7b</a></p><p>• How to build a company that withstands any era | Eric Ries, Lean Startup author: <a href="https://www.lennysnewsletter.com/p/how-to-build-a-company-that-withstands" target="_blank">https://www.lennysnewsletter.com/p/how-to-build-a-company-that-withstands</a></p><p>• <em>Fallout </em>on Prime Video: <a href="https://www.amazon.com/dp/B0CN4GGGQ2" target="_blank">https://www.amazon.com/dp/B0CN4GGGQ2</a></p><p>• <em>Fallout</em> (video game): <a href="https://fallout.bethesda.net" target="_blank">https://fallout.bethesda.net</a></p><p>• Claude Tag: <a href="https://www.anthropic.com/news/introducing-claude-tag" target="_blank">https://www.anthropic.com/news/introducing-claude-tag</a></p><p>—</p><p><strong>Recommended books:</strong></p><p>•<em> Crucial Conversations: Tools for Talking When Stakes Are High</em>: <a href="https://www.amazon.com/dp/0071771328" target="_blank">https://www.amazon.com/dp/0071771328</a></p><p>• <em>How to Raise an Adult: Break Free of the Overparenting Trap and Prepare Your Kid for Success</em>: <a href="https://www.amazon.com/How-Raise-Adult-Overparenting-Prepare/dp/1627791779" target="_blank">https://www.amazon.com/How-Raise-Adult-Overparenting-Prepare/dp/1627791779</a></p><p>• <em>Incorruptible: Why Good Companies Go Bad... and How Great Companies Stay Great</em>: <a href="https://www.amazon.com/dp/B0FWZZBPZB" target="_blank">https://www.amazon.com/dp/B0FWZZBPZB</a></p><p>—</p><p>Production and marketing by <a href="https://penname.co/" target="_blank">https://penname.co/</a>. For inquiries about sponsoring the podcast, email <a href="mailto:[email protected]" target="_blank">[email protected]</a>.</p><p>—</p><p><em>Lenny may be an investor in the companies discussed.</em></p> <br /><br />To hear more, visit <a href="https://www.lennysnewsletter.com?utm_medium=podcast&#38;utm_campaign=show-notes-no-free-preview-language">www.lennysnewsletter.com</a>

Key Insights

  • Diane Penn argues that in 2023 when she joined, nobody associated Anthropic and coding together, but recognizing long-form code generation as an emerging user behavior became an inflection point for the company's differentiation strategy.
  • Penn claims that 'evals are the new PRDs' at Anthropic—rather than writing traditional product requirements documents, the team now formalizes user pain points as comprehensive test cases with both failure and success scenarios that become the source of truth for model improvements.
  • Penn explains that emerging capabilities in larger models often appear discontinuously rather than smoothly, meaning teams may unknowingly train models with new abilities until evals reveal them, making continuous evaluation infrastructure critical to discovering what's possible.
  • Penn argues that product and model development are interdependent: Opus 4.5 wouldn't have succeeded without Cloud Code as a vehicle for the capability, and Cloud Code wouldn't have achieved adoption without Opus 4.5's intelligence level, creating a co-enabling relationship.
  • Penn states that the most creative thinkers at Anthropic spend substantial time directly working with Claude and research models because there's no substitute for hands-on experimentation when technology is moving this quickly—token spending is really an input to experimentation, not the goal itself.
  • Penn contends that having safety, alignment, and constitutional guardrails actually makes Claude a better conversational partner because it enables Claude to push back appropriately, rather than simply agreeing with users, which supports better thinking outcomes.
  • Penn claims that her role in product management for research involves translating vague user feedback ('Claude hallucinated') into specific, actionable problems researchers can understand and measure, such as distinguishing failures in tool use versus knowledge retrieval versus alignment.
  • Penn argues that even senior product managers and leaders at Anthropic must remain deeply hands-on with the technology, reading transcripts, building custom skills, and shipping work themselves—not just delegating—to maintain decision-making capability.
  • Penn explains that experimentation and discovery of AI capabilities is not an individual sport; her team found the most creative solutions emerged when employees shared ideas in shared channels and others iterated on those ideas, creating a 'mind meld' effect.
  • Penn asserts that the core value PMs bring is identifying what should be built and ensuring what was built is correct and good—this becomes harder, not easier, as models become more capable, making relentless user-centric work and first-principles thinking more essential.
  • Penn reveals that during periods of extreme shipping velocity (multiple model series per quarter), sustainability comes not from individuals managing heavy loads but from culture where low-ego team members actively support each other and protect colleagues on PTO by handling priorities.
  • Penn notes that at JP Morgan as a bond trader, she learned that the best ideas matter most regardless of the speaker's background or seniority, and she intentionally applies this by making herself vulnerable and authentic with her team to ensure diverse ideas surface.

Topics

Product management methodology evolution (evals replacing PRDs)Model training and capability emergence in frontier AIProduct-model co-development (product enablement of model capabilities)Anthropic's Labs innovation structureSkills and organizational culture at high-growth AI companiesWorking with AI as augmentation rather than replacementFrom coding as differentiator to agentic AI applicationsToken spending and 'living in the future' by 2028 standardsSafety, alignment, and constitution-based model behaviorPM roles in the AI era and hands-on technology requirementsExperimentation culture and communal discoveryBurnout prevention through team collaboration

Transcript

In 2023, when I started, nobody said anthropic and clod and coding in the same sentence. I want to go back to the beginning of anthropic. I remember feeling, man, these guys have no chance. OpenAI is so far ahead. At the time, I saw people were starting to use these models, not just for code autocomplete, but actually writing long-form code. Is that an opportunity for us to train Opus 3 to be better at. That was the inflection. I always think about Opus 4.5 a year later during winter break when everyone was home able to code. What was magical about Opus 4.5 is we also now not just had a model but a vehicle, a great product…

Full transcript available for MurmurCast members

Sign Up to Access

More from Lenny's Podcast: Product | Career | Growth

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.