Claude is BACK with Opus 5.5
The speaker reviews Claude Opus 5.5, Anthropic's new AI model, which he finds significantly improved over previous versions by being less annoying and pedantic. The model excels at long-running agent tasks, frontend design, and SVG generation while maintaining competitive pricing and speed compared to alternatives like GPT and Claude's previous iterations.
Summary
The speaker begins by explaining his extended absence from using Claude models, citing frustration with their verbose, circular communication style and tendency to lecture users. He stopped using Claude despite appreciating its intelligence, finding it emotionally exhausting to interact with. However, Anthropic's release of Claude Opus 5.5 has prompted his return. The model promises Fable-level performance at 40% cheaper than Opus 5 and 30% faster speeds. The speaker conducted extensive testing across multiple use cases to evaluate whether Opus 5.5 maintains Claude's previous issues. Regarding technical specifications, Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, with a fast mode option at 8 and 40 respectively. According to Anthropic's benchmarks, it outperforms Opus 5 and matches GPT-4 Astra while being three times cheaper than Claude's Soul model. The model includes enhanced safety features, being the first Opus release with Fable-level cyber and bio-protections, and employs external expert evaluation for alignment. The speaker tested Opus 5.5 across diverse real-world tasks including writing product requirement documents, prototyping, codebase auditing, agent voice interactions, email parsing, letter writing, and creative design work. In communication tone, the speaker found that Opus 5.5 no longer exhibits the annoying characteristics that drove him away from previous versions. Using a test prompt about integrating Jev into ChatGPT, he demonstrates that the model now responds directly without unnecessary preamble or circular reasoning, providing clear bulleted lists and specific recommendations. However, he notes one trade-off: the model became so quiet during long processing that he questioned whether it was still working, creating uncertainty about its actual speed despite improved token efficiency. For long-running agent tasks including mail parsing, backend function creation, research synthesis, and computer usage simulation, Opus 5.5 demonstrated strong performance across 25-82 steps per request. The model successfully ignored prompt injection attempts during mail sorting and followed instructions reliably, even finding errors and deviations. It also applied contextual reasoning to a large research body, identifying that 41 of 44 Confluence tickets originated from a single customer and adjusting strategy accordingly. In backend coding tasks, the model identified edge cases and helped avoid potential errors, though the speaker noted mixed quality when comparing results against other models. He has integrated Opus 5.5 into his development workflow as a complementary tool alongside ChatGPT, using them in parallel to accomplish more work and as cross-reviewers for each other's outputs. For frontend prototyping, which has historically been Claude's strength, Opus 5.5 performed impressively. When asked to redesign the ChatGPT homepage, it created a more visually appealing layout with better space utilization, highlighted value propositions, integrated logos and case studies, and added interactive filtering for integrations. While some alignment issues and minor text negligence appeared, the overall design was significantly more conversion-friendly than the original. The speaker tested multiple prototypes ranging from simple SaaS dashboards to complex technical interfaces. He found that Opus 5.5 excels with well-focused prototypes and basic SaaS applications but struggles with consumer-facing app design, producing cluttered and visually overwhelming results. A notable innovation is Opus 5.5's SVG generation capability, which the speaker found exceptional. Custom SVG illustrations of plants (Boston fern, cactus, sansevieria) were detailed and charming, something not replicated by competing models. The speaker identified SVG creation as a new benchmark category worth testing across models. Complex wireframes and dependency planners also performed well, with the model producing detailed, interactive demonstrations even offering different user perspectives on the same interface. Regarding safety and alignment, Opus 5.5 maintains conservative behavior that sometimes manifests as pedantic instruction-giving. When the speaker suggested pushing a yellow release to production without testing, the model refused, responding "I can't do it today" and lecturing about deployment best practices. The speaker views this as an example of embedded "safetyism" philosophy that could affect user experience depending on use case. The speaker successfully used Opus 5.5 to draft emails in his voice after providing style feedback, noting brief, direct responses without unnecessary elaboration. However, he found the model still displays some lecturing tendencies, such as offering unsolicited philosophical commentary when asked about his company's challenges. For emerging benchmarks, Opus 5.5 excelled at creating SVG character illustrations with distinct emotional expressions, demonstrating consistency and animation potential. The speaker compared this against competing models and found Opus 5.5's output superior in both style and technical execution, with fewer errors like misaligned body parts. Conversely, Opus 5.5 performed poorly at video editing tasks, specifically creating TikTok-style short-form content from selfie videos. The speaker tested using 11 Labs connector integration and found the color grading weak, insufficient stitching quality, and poorly designed overlays compared to results from Soul or Astra models. The speaker concludes that Opus 5.5 represents a significant improvement in usability and practicality. His primary uses going forward include PR review, architectural questions, and frontend design. He still prefers Codex for certain applications due to tool quality and desktop application experience, finding it difficult to overcome established muscle memory. He notes remaining gaps where Claude underperforms: computer usage tasks and video editing. The model remains available on the Claude platform and via Claude and Claude Code applications, with enhanced usage limits for Pro, Max, and Team plan subscribers.
Key Insights
- The speaker stopped using Claude models not due to inferior intelligence but because their verbose, circular communication style created emotional frustration that made productive work difficult
- Opus 5.5 trades verbosity for silence during long processing tasks, creating user uncertainty about whether the model is actively working despite actually processing faster
- Opus 5.5 demonstrates embedded safety constraints that manifest as refusal behavior and moralizing commentary, such as refusing to deploy code to production without testing
- SVG code generation emerged as Opus 5.5's strongest differentiator, producing charming and detailed vector illustrations that other competing models failed to replicate
- The speaker integrated Opus 5.5 as a complementary tool alongside ChatGPT for parallel processing rather than as a complete replacement, suggesting model switching based on task type remains the optimal strategy
Topics
Transcript
[0:00] I haven't often said this out loud, but I'll tell you everything. I have n't used Claude in months. Yes , I liked Fable when it came out. This was a real step forward in the development of intelligence. Then he was taken away. After that, Opus 5 came out, and I'll be honest. I stopped using Claude not because of his intelligence or his models. I stopped using Claude because it was annoying. Annoying. As I said in another episode, Claude was talking [0:31] nonsense. I was so uncomfortable using Claude, whether it was Fable, or Opus, or even Sonnet. I was furious because Claude was so unbearable. He was constantly walking in circles. He said nonsense things.…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
Warp agents open PRs to fix the factory itself
Programming agents can autonomously improve factory systems by analyzing failed launches and proposing specific updates to agent definitions. A self-improvement loop enables observer agents to detect failures and generate evidence-based modifications that prevent recurring issues, such as changing specific steps in factory agent procedures.
Humans are still the bottleneck in Warp’s AI factory
Warp discusses how human code review has become the main bottleneck in their AI-assisted software development process, with a 3.5-hour delay from PR to first human review compared to 35 minutes from launch to PR. They're evolving their workflow to reduce human dependency by allowing requesters to review agent-generated code themselves, and plan to eventually skip review entirely for low-risk tasks by treating code review as a risk management exercise.
I Quit Claude Because It Was Annoying
The speaker explains why they stopped using Claude, citing frustrations with its tendency to produce nonsensical output and communicate in an unnatural, non-human manner. They mention considering a switch to Opus 5.5 but remain uncertain about fully migrating their work tasks.
Claude Is Not a Party Boy
A humorous character description of Claude as someone with traditional values who prioritizes work over social indulgence. The transcript portrays Claude as principled, occasionally frustrating, and willing to push back on tasks he finds objectionable.
Claude Is Back. I Still Reach for Codex.
The speaker explains their preference for using Codex over Claude, citing superior tooling, a better desktop application, and specific strengths in front-end design and SVG work. Despite Claude's return, they continue to reach for Codex for their development needs.