The How I AI Bench
The speaker introduces 'The How I AI Bench,' a new set of human and AI-graded benchmarks designed to evaluate language models on practical tasks like writing PRDs, solving bugs, and designing systems. They test Claude Sonnet 3.5 against these benchmarks and note that while it scores lower than some specialized benchmarks (69% on Agentic Coding SweetBench Pro, 82% on Terminal Bench 2.1), the difference may not be noticeable in real-world usage.
Summary
The speaker expresses frustration with relying solely on subjective 'vibe checks' to evaluate AI models and announces the creation of 'The How I AI Bench'—a standardized set of benchmarks combining human and AI grading. The benchmark focuses on three practical competencies: writing PRDs (product requirement documents), solving bugs, and executing one-shot design tasks. These benchmarks are intended to be regularly used to assess new models as they release. The speaker then applies this benchmark to Claude Sonnet 3.5, presenting comparative performance data. While acknowledging that Sonnet 3.5 doesn't reach the top scores on other specialized benchmarks like Agentic Coding SweetBench Pro (69%) or Terminal Bench 2.1 (82%), the speaker suggests these marginal differences are unlikely to significantly impact most users' actual experience with the model. The speaker also notes the model's purported strengths in computer work and knowledge work tasks.
Key Insights
- The speaker is moving away from subjective 'vibe checks' toward developing standardized benchmarks that measure AI model performance on practical tasks users actually care about
- The How I AI Bench is specifically designed to test three practical competencies: writing PRDs, solving bugs, and executing one-shot design tasks using a combination of human and AI grading
- Claude Sonnet 3.5 scores 69% on Agentic Coding SweetBench Pro and 82% on Terminal Bench 2.1, positioning it as competitive but not leading on specialized benchmarks
- The speaker believes that the performance gaps between Sonnet 3.5 and higher-scoring models on technical benchmarks are unlikely to be noticed by most users in practical applications
- Claude Sonnet 3.5 is reported to have particular strengths in computer work and knowledge work tasks
Topics
Transcript
[0:00] I've been testing a lot of models and I'm starting to get bored of doing the vibe check. What I want to start developing is a set of benchmarks we can regularly test these new models against that you'll care about. So today I'm going to be introducing the how I AI bench, a set of AI and Clarvo graded benchmarks that are going to tell us if this model and any model is good at writing PRDs, solving bugs, and oneshotting designs. And we are going to [0:31] put Sonnet 5 to the test against that proposition. So, as you can see here, it's not quite at this [music] 69% on Agentic Coding SweetBench Pro or the 82%…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
The biggest barrier to AI adoption is not fear of technology. It's the absence of muscle memory.
The speaker argues that the primary barrier to AI adoption is not fear but lack of muscle memory—the habitual use of AI tools. Success requires building collaborative practices through empathy, accessible training, and establishing regular workflows rather than overcoming technological anxiety.
Staying in Gmail means none of your AI work compounds
The speaker discusses how staying within Gmail limits the compounding value of AI-assisted work and writing, and demonstrates a workflow using Claude to enhance Gmail's functionality by adding formatting, links, and draft management capabilities.
Intent engineering beats prompt engineering every time
The speaker emphasizes the importance of integrating AI, specifically Claude, into daily workflows by creating reminders to seek its assistance. This approach allows for collaborative problem-solving, transforming static prompts into dynamic conversations.
Claude Code for normal people: skills, voice mode, and how to collaborate with AI
The episode discusses how AI tools like Claude can enhance productivity and collaboration in business and personal tasks by enabling users to focus on intentions rather than complex prompting. Grace Clark offers insights on her projects and how she built tailored solutions to improve communication and client relations using AI.
Running evals on your internal agents is how you keep them trustworthy over time
The evaluation of internal AI agents is crucial for maintaining their reliability. An internal eval platform allows engineers to review agent performance regularly and ensure their responses are accurate, particularly in critical areas like coding.