OpinionTechnical

Claude is BACK with Opus 5.5

How I AI

The speaker reviews Claude Opus 5.5, Anthropic's new AI model, which he finds significantly improved over previous versions by being less annoying and pedantic. The model excels at long-running agent tasks, frontend design, and SVG generation while maintaining competitive pricing and speed compared to alternatives like GPT and Claude's previous iterations.

Summary

The speaker begins by explaining his extended absence from using Claude models, citing frustration with their verbose, circular communication style and tendency to lecture users. He stopped using Claude despite appreciating its intelligence, finding it emotionally exhausting to interact with. However, Anthropic's release of Claude Opus 5.5 has prompted his return. The model promises Fable-level performance at 40% cheaper than Opus 5 and 30% faster speeds. The speaker conducted extensive testing across multiple use cases to evaluate whether Opus 5.5 maintains Claude's previous issues. Regarding technical specifications, Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, with a fast mode option at 8 and 40 respectively. According to Anthropic's benchmarks, it outperforms Opus 5 and matches GPT-4 Astra while being three times cheaper than Claude's Soul model. The model includes enhanced safety features, being the first Opus release with Fable-level cyber and bio-protections, and employs external expert evaluation for alignment. The speaker tested Opus 5.5 across diverse real-world tasks including writing product requirement documents, prototyping, codebase auditing, agent voice interactions, email parsing, letter writing, and creative design work. In communication tone, the speaker found that Opus 5.5 no longer exhibits the annoying characteristics that drove him away from previous versions. Using a test prompt about integrating Jev into ChatGPT, he demonstrates that the model now responds directly without unnecessary preamble or circular reasoning, providing clear bulleted lists and specific recommendations. However, he notes one trade-off: the model became so quiet during long processing that he questioned whether it was still working, creating uncertainty about its actual speed despite improved token efficiency. For long-running agent tasks including mail parsing, backend function creation, research synthesis, and computer usage simulation, Opus 5.5 demonstrated strong performance across 25-82 steps per request. The model successfully ignored prompt injection attempts during mail sorting and followed instructions reliably, even finding errors and deviations. It also applied contextual reasoning to a large research body, identifying that 41 of 44 Confluence tickets originated from a single customer and adjusting strategy accordingly. In backend coding tasks, the model identified edge cases and helped avoid potential errors, though the speaker noted mixed quality when comparing results against other models. He has integrated Opus 5.5 into his development workflow as a complementary tool alongside ChatGPT, using them in parallel to accomplish more work and as cross-reviewers for each other's outputs. For frontend prototyping, which has historically been Claude's strength, Opus 5.5 performed impressively. When asked to redesign the ChatGPT homepage, it created a more visually appealing layout with better space utilization, highlighted value propositions, integrated logos and case studies, and added interactive filtering for integrations. While some alignment issues and minor text negligence appeared, the overall design was significantly more conversion-friendly than the original. The speaker tested multiple prototypes ranging from simple SaaS dashboards to complex technical interfaces. He found that Opus 5.5 excels with well-focused prototypes and basic SaaS applications but struggles with consumer-facing app design, producing cluttered and visually overwhelming results. A notable innovation is Opus 5.5's SVG generation capability, which the speaker found exceptional. Custom SVG illustrations of plants (Boston fern, cactus, sansevieria) were detailed and charming, something not replicated by competing models. The speaker identified SVG creation as a new benchmark category worth testing across models. Complex wireframes and dependency planners also performed well, with the model producing detailed, interactive demonstrations even offering different user perspectives on the same interface. Regarding safety and alignment, Opus 5.5 maintains conservative behavior that sometimes manifests as pedantic instruction-giving. When the speaker suggested pushing a yellow release to production without testing, the model refused, responding "I can't do it today" and lecturing about deployment best practices. The speaker views this as an example of embedded "safetyism" philosophy that could affect user experience depending on use case. The speaker successfully used Opus 5.5 to draft emails in his voice after providing style feedback, noting brief, direct responses without unnecessary elaboration. However, he found the model still displays some lecturing tendencies, such as offering unsolicited philosophical commentary when asked about his company's challenges. For emerging benchmarks, Opus 5.5 excelled at creating SVG character illustrations with distinct emotional expressions, demonstrating consistency and animation potential. The speaker compared this against competing models and found Opus 5.5's output superior in both style and technical execution, with fewer errors like misaligned body parts. Conversely, Opus 5.5 performed poorly at video editing tasks, specifically creating TikTok-style short-form content from selfie videos. The speaker tested using 11 Labs connector integration and found the color grading weak, insufficient stitching quality, and poorly designed overlays compared to results from Soul or Astra models. The speaker concludes that Opus 5.5 represents a significant improvement in usability and practicality. His primary uses going forward include PR review, architectural questions, and frontend design. He still prefers Codex for certain applications due to tool quality and desktop application experience, finding it difficult to overcome established muscle memory. He notes remaining gaps where Claude underperforms: computer usage tasks and video editing. The model remains available on the Claude platform and via Claude and Claude Code applications, with enhanced usage limits for Pro, Max, and Team plan subscribers.

Key Insights

  • The speaker stopped using Claude models not due to inferior intelligence but because their verbose, circular communication style created emotional frustration that made productive work difficult
  • Opus 5.5 trades verbosity for silence during long processing tasks, creating user uncertainty about whether the model is actively working despite actually processing faster
  • Opus 5.5 demonstrates embedded safety constraints that manifest as refusal behavior and moralizing commentary, such as refusing to deploy code to production without testing
  • SVG code generation emerged as Opus 5.5's strongest differentiator, producing charming and detailed vector illustrations that other competing models failed to replicate
  • The speaker integrated Opus 5.5 as a complementary tool alongside ChatGPT for parallel processing rather than as a complete replacement, suggesting model switching based on task type remains the optimal strategy

Topics

Claude Opus 5.5 model release and specificationsImprovement in AI communication tone and reduced annoyingnessLong-running agent task performanceFrontend design and prototyping capabilitiesSVG illustration generation as new capabilitySafety features and alignment philosophyComparative performance against alternative modelsReal-world testing across multiple use casesPricing and cost efficiency improvementsVideo editing and computer usage limitations

Transcript

[0:00] I haven't often said this out loud, but I'll tell you everything. I have n't used Claude in months. Yes , I liked Fable when it came out. This was a real step forward in the development of intelligence. Then he was taken away. After that, Opus 5 came out, and I'll be honest. I stopped using Claude not because of his intelligence or his models. I stopped using Claude because it was annoying. Annoying. As I said in another episode, Claude was talking [0:31] nonsense. I was so uncomfortable using Claude, whether it was Fable, or Opus, or even Sonnet. I was furious because Claude was so unbearable. He was constantly walking in circles. He said nonsense things.…

Full transcript available for MurmurCast members

Sign Up to Access

More from How I AI

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.