Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn
Diane Penn, Anthropic's first technical PM, discusses how the company evolved from an underdog to a $50B+ ARR leader by combining frontier models with exceptional products. She explains how PMs now use 'evals as PRDs,' prioritize experimentation over perfect planning, and must stay deeply hands-on with AI technology to guide teams effectively.
Summary
Diane Penn joined Anthropic in 2023 as its first technical product manager when the company had only five engineers and no released model. Despite initial skepticism that OpenAI had already won the race, Anthropic found its identity through strategic focus on coding as a differentiating use case. Penn traces key inflection points: Opus 3's launch (early 2024) proved Anthropic could build frontier models; the Golden Gate Bridge feature showcased how research could be brought to users quickly; and Opus 4.5 (2024) demonstrated the power of pairing advanced models with exceptional products like Cloud Code.
Penn describes a fundamental shift in product management methodology at Anthropic. Rather than traditional PRDs, the team now uses 'evals as the new PRDs'—they identify specific user pain points through transcripts and feedback, write comprehensive test cases that capture both failure scenarios and success criteria, and track model improvements against these evals. This emerged from concrete problems: early users reported Claude couldn't follow schemas properly, so the team created 40 examples of JSON formatting failures, formalized them as an eval, and measured progress across model versions until the problem was solved to near-perfection.
The company operates in what Penn calls 'the exponential,' where each model generation brings discontinuous capability jumps. Rather than predict exactly what's possible, the organization maintains adaptability through continuous evals and experimentation. Labs functions as an internal innovation engine that identifies 'discontinuous large bets' outside the core roadmap—including Cloud Code, Skills, Cloud Design, and MCP—using small, self-driven teams with strong opinions on themes but flexible approaches to prototypes.
Penn emphasizes that product management has fundamentally changed in the AI era. Managers must spend significant time hands-on with the technology, 'sweating tokens as much as pixels.' She uses Claude extensively herself—building custom Skills to improve her management capabilities, consulting it on pricing decisions, and using it to prepare for crucial conversations. She maintains her own point of view by thinking first, then using Claude as a sparring partner that pushes back rather than simply agrees.
The company's culture prioritizes collaboration and experimentation over individual work. Penn describes discovering use cases in shared Slack channels where employees test new versions and iterate rapidly on ideas together. She attributes Anthropic's ability to ship multiple model series per quarter to radical ownership, team collaboration, and deliberate hiring for low-ego, mission-aligned individuals who support each other. She explicitly rejects the notion that PMs are unnecessary when models are powerful; instead, the relentless work of understanding user needs, providing actionable feedback to researchers, and discovering what's possible with new capabilities makes PMs more essential than ever.
About this episode
<p><strong>Dianne Penn</strong> is Head of Product for Anthropic’s AI Research and Labs teams. She joined in 2023 as Anthropic’s first technical product manager, when the entire product team was five engineers, and has since helped ship every model from Claude 2 through Fable, and helped incubate Claude Code, MCP, Skills, computer use, tool use, and reasoning. Before Anthropic, she helped build Alexa’s AI at Amazon and, before that, traded high-yield bonds at JP Morgan Chase.</p><p></p><p><strong>In our in-depth conversation, we discuss:</strong></p><p>1. What Anthropic’s early days were like</p><p>2. The inflection points that turned Anthropic from an underdog into the fastest-growing company in history</p><p>3. How exactly Claude got so good at coding</p><p>4. The eval-driven development loop her team is pioneering</p><p>5. How to find joy in AI when everything is moving this fast</p><p>6. Why Claude’s willingness to push back is key to its success</p><p>7. Where human judgment remains irreplaceable</p><p>—</p><p><strong>Brought to you by:</strong></p><p><a href="https://workos.com/lenny" target="_blank"><strong>WorkOS</strong></a>—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more</p><p><a href="https://mercury.com/command?utm_source=lennys&utm_medium=sponsored_newsletter&utm_campaign=26q3_brand_campaign" target="_blank"><strong>Mercury</strong></a>—Radically different banking, now with Command</p><p>—</p><p><strong>Episode transcript: </strong><a href="https://www.lennysnewsletter.com/p/anthropics-first-technical-pm-on" target="_blank">https://www.lennysnewsletter.com/p/anthropics-first-technical-pm-on</a></p><p>—</p><p><strong>Archive of all Lenny's Podcast transcripts: </strong><a href="https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&st=ahz0fj11&dl=0" target="_blank">https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&st=ahz0fj11&dl=0</a></p><p>—</p><p><strong>Where to find Dianne Penn:</strong></p><p>• LinkedIn: <a href="http://linkedin.com/in/dianne-na-penn" target="_blank">linkedin.com/in/dianne-na-penn</a></p><p>—</p><p><strong>Where to find Lenny:</strong></p><p>• Newsletter: <a href="https://www.lennysnewsletter.com" target="_blank">https://www.lennysnewsletter.com</a></p><p>• X: <a href="https://twitter.com/lennysan" target="_blank">https://twitter.com/lennysan</a></p><p>• LinkedIn: <a href="https://www.linkedin.com/in/lennyrachitsky/" target="_blank">https://www.linkedin.com/in/lennyrachitsky/</a></p><p>—</p><p><strong>In this episode, we cover:</strong></p><p>(00:00) Introduction</p><p>(02:31) Early Anthropic days</p><p>(08:55) Big milestones</p><p>(13:50) Inside the exponential</p><p>(20:02) Token maxing</p><p>(23:30) Anthropic Labs and the incubation model</p><p>(27:30) How the research role works</p><p>(31:35) How to become a top researcher</p><p>(35:18) Frontier model safeguards</p><p>(39:38) Hiring in the AI era</p><p>(44:16) Building an eval set</p><p>(47:48) Evals vs PRDs</p><p>(49:55) The importance of hands-on leadership</p><p>(52:46) Finding joy in AI</p><p>(58:10) How Dianne uses Claude</p><p>(01:01:05) Avoiding overreliance on AI</p><p>(01:03:50) The constitution that makes Claude better</p><p>(01:07:11) AI writing and verification</p><p>(01:11:40) Where human brains will continue to be valuable</p><p>(01:14:10) Navigating AI with kids</p><p>(01:16:26) Alignment, the future of the PM role, and burnout</p><p>(01:21:54) Lightning round and final thoughts</p><p>—</p><p><strong>Referenced:</strong></p><p>• Anthropic: <a href="https://www.anthropic.com" target="_blank">https://www.anthropic.com</a></p><p>• Golden Gate Claude: <a href="https://www.anthropic.com/news/golden-gate-claude" target="_blank">https://www.anthropic.com/news/golden-gate-claude</a></p><p>• Dario Amodei’s website: <a href="https://darioamodei.com" target="_blank">https://darioamodei.com</a></p><p>• Scaling Laws and Interpretability of Learning from Repeated Data: <a href="https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data" target="_blank">https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data</a></p><p>• Tokenmaxxing: How Top Builders Use AI To Do The Work Of 400 Engineers: <a href="https://www.ycombinator.com/library/Pa-tokenmaxxing-how-top-builders-use-ai-to-do-the-work-of-400-engineers" target="_blank">https://www.ycombinator.com/library/Pa-tokenmaxxing-how-top-builders-use-ai-to-do-the-work-of-400-engineers</a></p><p>• Garry Tan on X: <a href="https://x.com/garrytan" target="_blank">https://x.com/garrytan</a></p><p>• Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night | Ben Mann: <a href="https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann" target="_blank">https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann</a></p><p>• Anthropic’s CPO on what comes next | Mike Krieger (co-founder of Instagram): <a href="https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next" target="_blank">https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next</a></p><p>• Introducing Labs: <a href="https://www.anthropic.com/news/introducing-anthropic-labs" target="_blank">https://www.anthropic.com/news/introducing-anthropic-labs</a></p><p>• Louis CK | about airplane Wi Fi: <a href="https://www.youtube.com/watch?v=me4BZBsHwZs" target="_blank">https://www.youtube.com/watch?v=me4BZBsHwZs</a></p><p>• What happens after coding is solved? | Fiona Fung (Manager of the Claude Code and Cowork Teams): <a href="https://www.lennysnewsletter.com/p/building-the-most-ai-pilled-engineering" target="_blank">https://www.lennysnewsletter.com/p/building-the-most-ai-pilled-engineering</a></p><p>• The Anthropic Hive Mind: <a href="https://steve-yegge.medium.com/the-anthropic-hive-mind-d01f768f3d7b" target="_blank">https://steve-yegge.medium.com/the-anthropic-hive-mind-d01f768f3d7b</a></p><p>• How to build a company that withstands any era | Eric Ries, Lean Startup author: <a href="https://www.lennysnewsletter.com/p/how-to-build-a-company-that-withstands" target="_blank">https://www.lennysnewsletter.com/p/how-to-build-a-company-that-withstands</a></p><p>• <em>Fallout </em>on Prime Video: <a href="https://www.amazon.com/dp/B0CN4GGGQ2" target="_blank">https://www.amazon.com/dp/B0CN4GGGQ2</a></p><p>• <em>Fallout</em> (video game): <a href="https://fallout.bethesda.net" target="_blank">https://fallout.bethesda.net</a></p><p>• Claude Tag: <a href="https://www.anthropic.com/news/introducing-claude-tag" target="_blank">https://www.anthropic.com/news/introducing-claude-tag</a></p><p>—</p><p><strong>Recommended books:</strong></p><p>•<em> Crucial Conversations: Tools for Talking When Stakes Are High</em>: <a href="https://www.amazon.com/dp/0071771328" target="_blank">https://www.amazon.com/dp/0071771328</a></p><p>• <em>How to Raise an Adult: Break Free of the Overparenting Trap and Prepare Your Kid for Success</em>: <a href="https://www.amazon.com/How-Raise-Adult-Overparenting-Prepare/dp/1627791779" target="_blank">https://www.amazon.com/How-Raise-Adult-Overparenting-Prepare/dp/1627791779</a></p><p>• <em>Incorruptible: Why Good Companies Go Bad... and How Great Companies Stay Great</em>: <a href="https://www.amazon.com/dp/B0FWZZBPZB" target="_blank">https://www.amazon.com/dp/B0FWZZBPZB</a></p><p>—</p><p>Production and marketing by <a href="https://penname.co/" target="_blank">https://penname.co/</a>. For inquiries about sponsoring the podcast, email <a href="mailto:[email protected]" target="_blank">[email protected]</a>.</p><p>—</p><p><em>Lenny may be an investor in the companies discussed.</em></p> <br /><br />To hear more, visit <a href="https://www.lennysnewsletter.com?utm_medium=podcast&utm_campaign=show-notes-no-free-preview-language">www.lennysnewsletter.com</a>
Key Insights
- Diane Penn argues that in 2023 when she joined, nobody associated Anthropic and coding together, but recognizing long-form code generation as an emerging user behavior became an inflection point for the company's differentiation strategy.
- Penn claims that 'evals are the new PRDs' at Anthropic—rather than writing traditional product requirements documents, the team now formalizes user pain points as comprehensive test cases with both failure and success scenarios that become the source of truth for model improvements.
- Penn explains that emerging capabilities in larger models often appear discontinuously rather than smoothly, meaning teams may unknowingly train models with new abilities until evals reveal them, making continuous evaluation infrastructure critical to discovering what's possible.
- Penn argues that product and model development are interdependent: Opus 4.5 wouldn't have succeeded without Cloud Code as a vehicle for the capability, and Cloud Code wouldn't have achieved adoption without Opus 4.5's intelligence level, creating a co-enabling relationship.
- Penn states that the most creative thinkers at Anthropic spend substantial time directly working with Claude and research models because there's no substitute for hands-on experimentation when technology is moving this quickly—token spending is really an input to experimentation, not the goal itself.
- Penn contends that having safety, alignment, and constitutional guardrails actually makes Claude a better conversational partner because it enables Claude to push back appropriately, rather than simply agreeing with users, which supports better thinking outcomes.
- Penn claims that her role in product management for research involves translating vague user feedback ('Claude hallucinated') into specific, actionable problems researchers can understand and measure, such as distinguishing failures in tool use versus knowledge retrieval versus alignment.
- Penn argues that even senior product managers and leaders at Anthropic must remain deeply hands-on with the technology, reading transcripts, building custom skills, and shipping work themselves—not just delegating—to maintain decision-making capability.
- Penn explains that experimentation and discovery of AI capabilities is not an individual sport; her team found the most creative solutions emerged when employees shared ideas in shared channels and others iterated on those ideas, creating a 'mind meld' effect.
- Penn asserts that the core value PMs bring is identifying what should be built and ensuring what was built is correct and good—this becomes harder, not easier, as models become more capable, making relentless user-centric work and first-principles thinking more essential.
- Penn reveals that during periods of extreme shipping velocity (multiple model series per quarter), sustainability comes not from individuals managing heavy loads but from culture where low-ego team members actively support each other and protect colleagues on PTO by handling priorities.
- Penn notes that at JP Morgan as a bond trader, she learned that the best ideas matter most regardless of the speaker's background or seniority, and she intentionally applies this by making herself vulnerable and authentic with her team to ensure diverse ideas surface.
Topics
Transcript
In 2023, when I started, nobody said anthropic and clod and coding in the same sentence. I want to go back to the beginning of anthropic. I remember feeling, man, these guys have no chance. OpenAI is so far ahead. At the time, I saw people were starting to use these models, not just for code autocomplete, but actually writing long-form code. Is that an opportunity for us to train Opus 3 to be better at. That was the inflection. I always think about Opus 4.5 a year later during winter break when everyone was home able to code. What was magical about Opus 4.5 is we also now not just had a model but a vehicle, a great product…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Lenny's Podcast: Product | Career | Growth
How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)
Roman Ugarte, product lead at Grokbot (SpaceX AI), discusses how they built and launched a groundbreaking AI agent product in just one month by making two critical early decisions: running bots in the cloud with their own computers, and focusing ruthlessly on simplicity over features. The conversation covers their go-to-market strategy, the importance of manual onboarding, and how maintaining startup culture enables rapid iteration in a highly competitive AI market.
Why companies are becoming a series of loops | Anish Acharya (a16z)
Anish Acharya, a16z General Partner and former founder, discusses how AI is reshaping company organization through 'loops' (autonomous agent workflows), argues the permanent underclass fear is overblown, and contends that the real consumer opportunity lies in AI applications that improve happiness and human connection rather than just productivity.
AI’s third era: the rise of persistent AI coworkers | Tara Seshan (OpenAI’s product lead)
Tara Seshan, OpenAI's product lead for ChatGPT and Codex, discusses how AI is fundamentally changing product management and knowledge work. She emphasizes the shift from theoretical planning to rapid empirical iteration, the importance of maintaining human ambition and opinionation in an age of AI capabilities, and how products like work mode and sites are enabling persistent AI coworkers.
How to close $100K+ enterprise deals, step by step | Jen Abel
Jen Abel breaks down the enterprise sales cycle into 15+ steps (not the commonly believed 5), emphasizing the importance of building relationships, collecting competitive intelligence, and maintaining tight project management throughout. She argues that successful enterprise sales requires listening more than pitching, understanding the buyer's internal dynamics, and creating a sense of co-authorship rather than traditional selling.
OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber
Ian Silber, Head of Design at OpenAI, discusses why this is the best time in history to be a designer despite widespread anxiety in the design community. He argues that while designers face uncertainty about changing roles, those who embrace AI tools as exploratory amplifiers rather than threats are experiencing unprecedented creative opportunities and agency.