The How I AI Bench
The speaker introduces 'The How I AI Bench,' a new set of human and AI-graded benchmarks designed to evaluate language models on practical tasks like writing PRDs, solving bugs, and designing systems. They test Claude Sonnet 3.5 against these benchmarks and note that while it scores lower than some specialized benchmarks (69% on Agentic Coding SweetBench Pro, 82% on Terminal Bench 2.1), the difference may not be noticeable in real-world usage.
Summary
The speaker expresses frustration with relying solely on subjective 'vibe checks' to evaluate AI models and announces the creation of 'The How I AI Bench'—a standardized set of benchmarks combining human and AI grading. The benchmark focuses on three practical competencies: writing PRDs (product requirement documents), solving bugs, and executing one-shot design tasks. These benchmarks are intended to be regularly used to assess new models as they release. The speaker then applies this benchmark to Claude Sonnet 3.5, presenting comparative performance data. While acknowledging that Sonnet 3.5 doesn't reach the top scores on other specialized benchmarks like Agentic Coding SweetBench Pro (69%) or Terminal Bench 2.1 (82%), the speaker suggests these marginal differences are unlikely to significantly impact most users' actual experience with the model. The speaker also notes the model's purported strengths in computer work and knowledge work tasks.
Key Insights
- The speaker is moving away from subjective 'vibe checks' toward developing standardized benchmarks that measure AI model performance on practical tasks users actually care about
- The How I AI Bench is specifically designed to test three practical competencies: writing PRDs, solving bugs, and executing one-shot design tasks using a combination of human and AI grading
- Claude Sonnet 3.5 scores 69% on Agentic Coding SweetBench Pro and 82% on Terminal Bench 2.1, positioning it as competitive but not leading on specialized benchmarks
- The speaker believes that the performance gaps between Sonnet 3.5 and higher-scoring models on technical benchmarks are unlikely to be noticed by most users in practical applications
- Claude Sonnet 3.5 is reported to have particular strengths in computer work and knowledge work tasks
Topics
Transcript
[0:00] I've been testing a lot of models and I'm starting to get bored of doing the vibe check. What I want to start developing is a set of benchmarks we can regularly test these new models against that you'll care about. So today I'm going to be introducing the how I AI bench, a set of AI and Clarvo graded benchmarks that are going to tell us if this model and any model is good at writing PRDs, solving bugs, and oneshotting designs. And we are going to [0:31] put Sonnet 5 to the test against that proposition. So, as you can see here, it's not quite at this [music] 69% on Agentic Coding SweetBench Pro or the 82%…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
Jev: 8 real use cases this fast, cheap model
Claire Val and returning guest John Lindquist discuss Jev, a fast and cost-effective decision model from Type Safe AI, exploring eight real-world use cases including task management, data reconciliation, real-time routing, and agent coordination. They contrast Jev's structured decision-making approach with traditional LLMs, emphasizing how its speed and affordability unlock previously impractical applications.
Jev analyzed 1,700 PRs for 9 cents
A developer used AI (Gemini) to analyze 1,700 pull requests in 2 minutes for just 9 cents, extracting work allocation data across initiatives. This demonstrates how AI can help CTOs and CEOs quantify what percentage of engineering effort goes toward different products or projects.
Jev clusters your data for precise AI actions
Jev is a tool that enables precise AI actions by allowing users to organize large bodies of information through tagging, categorization, clustering, and filtering. The system applies targeted AI operations to specific data clusters, with practical applications including error severity sorting and intelligent email processing.
I tried Muse, Meta's new AI agent (meet Slime 🦖)
The speaker reviews Muse, Meta's new AI agent, highlighting its ability to perform browser-based tasks, connect to multiple data sources, and help users achieve personal goals through an approachable interface. Key features include connectors to email/calendar/health data, idea suggestions, artifact creation, and customizable avatars.
I’m using Jev more than Opus 5.5 or GPT-6. Here’s why.
The speaker demonstrates why they're using Jev, a fast and inexpensive decision-making model from Type-Safe AI, more than other recent models like Opus 5.5 and GPT-6. Jev specializes in classification, clustering, and real-time decision-making tasks at a fraction of the cost of traditional LLMs, enabling complex data analysis and product features that would have been prohibitively expensive before.