I’m using Jev more than Opus 5.5 or GPT-6. Here’s why.
The speaker demonstrates why they're using Jev, a fast and inexpensive decision-making model from Type-Safe AI, more than other recent models like Opus 5.5 and GPT-6. Jev specializes in classification, clustering, and real-time decision-making tasks at a fraction of the cost of traditional LLMs, enabling complex data analysis and product features that would have been prohibitively expensive before.
Summary
The speaker opens by noting the recent release of multiple major AI models but argues that Jev, a decision-making model from Type-Safe AI, has proven more useful for their work across personal productivity, coding, and product development. Jev differs fundamentally from standard LLMs in that it accepts unstructured text input but returns type-safe values—predefined options like choices, ratings, or boolean probability scores—rather than generated text. This makes it exceptionally cheap (4 cents per million input tokens, with no output token charges) and fast, making it suitable for real-time applications.
The speaker demonstrates three major use cases: First, they used Jev to analyze thousands of GitHub pull requests by having it classify and cluster PRs to understand work distribution across their projects. On a marketing site with 112 PRs, this cost 1.1 cents and took two minutes; on a production app with 2,000 PRs, it cost nine cents total. Second, they analyzed their local Claude/Codex sessions to understand how their time allocation has shifted from pure engineering to agent work and media production. Third, they ran Jev on their personal Gmail to classify emails for deletion.
The speaker's most ambitious project uses Jev as part of a hybrid architecture: collecting 1,000-1,100 signals from PRs, support tickets, conversations, and Linear tickets, using Jev for fast classification and clustering, then applying more sophisticated models (Astra) for deeper analysis. This produced over 200,000 classifications and pairwise groupings for approximately $4 in Jev costs, unlocking a product insights graph that shows gaps between customer requests and actual development work.
The speaker then demonstrates two real-time applications built in one evening: First, a YouTube comment analyzer that pulls 4,500 comments from the 'How I Build AI' channel, classifies them as positive/negative/neutral, identifies episode suggestions, and provides searchable dashboards. Second, a real-time emotion-to-quote app that uses OpenAI's voice API with Jev to determine emotional state, select matching colors from a predefined palette, and retrieve appropriate quotes—demonstrating how Jev enables instant decision-making in interactive applications.
Key Insights
- Jev charges only for input tokens at 4 cents per million, with no output token fees, making it drastically cheaper than standard LLMs for classification tasks
- The speaker analyzed 2,000 production PRs to determine work distribution across initiatives, found 17,000 matching pairs, and tagged themes for $0.09 total—a task they claim would have cost $100,000 three years ago
- Jev excels at pairwise comparison tasks, determining whether two items are related by answering yes/no questions, which enables clustering of large datasets without the model needing to generate explanations
- The most complex project collected 1,100 raw signals from multiple business sources and generated over 200,000 classifications using Jev for $4, creating an insights graph that revealed gaps between customer requests and development priorities
- Jev's ability to make real-time decisions from predefined options enables interactive applications like emotion detection with instant quote matching, which would have unacceptable latency with traditional generative models
Topics
Transcript
[0:00] Jeev, Jeev, Jeev. Welcome to Jev's week on the How I Build AI channel. Over the past 5 days, we have seen the release of many new models. We've seen Opus 5.5, we've seen GPT-6 Soul, GPT- 6 Luna, Muse is blowing up the newsfeed. Everyone still loves their Grok bots. And yet, there's one thing I want to talk about in the field of AI right now. [0:31] This is a fast, cheap, non-talking, first-system decision-making model from Type-Safe AI. As soon as I saw it in X trends, as soon as I saw the launch, I immediately started testing it. I have to say that more than any other model I've tried recently, Jev has been the…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
Warp agents open PRs to fix the factory itself
Programming agents can autonomously improve factory systems by analyzing failed launches and proposing specific updates to agent definitions. A self-improvement loop enables observer agents to detect failures and generate evidence-based modifications that prevent recurring issues, such as changing specific steps in factory agent procedures.
Humans are still the bottleneck in Warp’s AI factory
Warp discusses how human code review has become the main bottleneck in their AI-assisted software development process, with a 3.5-hour delay from PR to first human review compared to 35 minutes from launch to PR. They're evolving their workflow to reduce human dependency by allowing requesters to review agent-generated code themselves, and plan to eventually skip review entirely for low-risk tasks by treating code review as a risk management exercise.
I Quit Claude Because It Was Annoying
The speaker explains why they stopped using Claude, citing frustrations with its tendency to produce nonsensical output and communicate in an unnatural, non-human manner. They mention considering a switch to Opus 5.5 but remain uncertain about fully migrating their work tasks.
Claude Is Not a Party Boy
A humorous character description of Claude as someone with traditional values who prioritizes work over social indulgence. The transcript portrays Claude as principled, occasionally frustrating, and willing to push back on tasks he finds objectionable.
Claude Was Fast. It Didn’t Feel Fast.
A developer discusses how Claude's actual speed didn't translate to perceived speed during long conversations because the model remained silent without vocalizing its actions, creating user uncertainty. The speaker emphasizes that real-world model performance, perceived latency, response verbosity, and response format all significantly impact end-user experience in agent development.