I reviewed Opus 5.5 and GPT-6 Sol live - and the results surprised me
A live review comparing three newly released AI models—Opus 5.5, GPT-6 Soul, and GPT-6 Luna—where the reviewer conducts blind testing across multiple task categories and finds that while Opus 5.5 significantly improves on previous versions, OpenAI's models excel in creative tasks like SVG generation.
Summary
The reviewer conducted a comprehensive live evaluation of three new AI models released the same morning: Claude Opus 5.5 by Anthropic, GPT-6 Soul, and GPT-6 Luna by OpenAI. The reviewer had early access to Opus 5.5 and expanded their 'How I AI' benchmark to test models across diverse categories including PRDs, personal productivity, frontend coding, backend coding, agent personality, long-term research, computer usage, and creative tasks like SVG illustration and video editing.
Key findings about pricing and performance: Opus 5.5 costs nearly twice as much as GPT-6 Soul, though both are cheaper than their predecessors (Opus 5 and Fable 1). All three models showed improvements in cost efficiency, speed, and token optimization through cached inputs. Anthropic implemented Fable-level security restrictions in Opus 5.5 for cybersecurity and biology tasks.
Regarding user experience and communication style: The reviewer noted significant improvements in Opus 5.5's interaction quality compared to Opus 5, describing previous versions as 'annoying' and 'conservative.' Opus 5.5 produces clearer bulleted lists and normal behavior, though it remains somewhat verbose and provides less commentary than OpenAI models. GPT-6 Soul was found to be faster in actual latency and provided better balance in communication level.
Testing results showed distinct model strengths: For PRD writing, Soul performed best. For agent personality and long-term tasks, Claude Opus 5.5 excelled. For personal productivity tasks, Model B and Model E (which appeared to be different OpenAI variants) wrote the most readable messages. Surprisingly, OpenAI models significantly outperformed Claude on SVG character creation, producing cuter and more detailed illustrations. For frontend design, Claude models tended to over-complicate interfaces while OpenAI models provided cleaner, more usable designs. All models performed poorly on video editing tasks.
The reviewer's overall assessment: Opus 5.5 won 'the week' by receiving the highest average scores across individual tasks, particularly excelling in agent voice work and specific design tasks. However, the reviewer's personal preference ('my heart') went to OpenAI's Astra and Soul models, particularly due to Soul's cheaper pricing and superior performance on creative tasks. Notably, an LLM judge disagreed with the reviewer's assessments, preferring Fable and ranking Soul lower, highlighting differences between human and AI-based evaluation criteria.
Key Insights
- Opus 5.5 represents a significant improvement in user experience and interaction quality compared to Opus 5, to the point where the reviewer stopped feeling irritated during everyday use and returned it to their daily workflow after months of preferring other models.
- Anthropic's design choices to make Opus 5.5 less verbose by reducing commentary about its work creates an unintended user experience issue where people question whether the model is actually working, making it feel slower than it actually is.
- OpenAI models (Astra and Soul) significantly outperform Claude models at SVG character creation and creative illustration tasks, contradicting the reviewer's initial prediction that Claude would excel at visual creative work.
- Claude models tend to over-complicate frontend designs with excessive detail and richness when given complex requests, making prototypes difficult to understand, whereas OpenAI models maintain cleaner, more readable interfaces.
- An LLM judge evaluating the reviewer's own test results disagreed substantially with their assessments, preferring Fable over Astra and rating Soul much lower, revealing fundamental differences in how AI judges versus humans evaluate model outputs.
Topics
Transcript
[0:00] I was as prepared as possible for Anthropic to release Opus 5.5. I was even prepared for OpenAI to release another model. I got up early this morning because I had early access to Opus 5.5 to record an incredible review for you . And guess what? They both, they both came out this morning. So now, even though I already have a great Opus 5.5 review ready, [0:30] which you'll see later, I'm just going to do this live stream. And I've never done live broadcasts before. So I'm going to tell you about these new models. Three models were released today: Opus 5.5 by Anthropic, GPT-6 Soul, and GPT-6 Luna. There are price wars going on here,…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
Barbie Bench: AGI Has Not Arrived
A content creator demonstrates Claude Opus's Barbie fashion designer 3D rendering project, highlighting both impressive and flawed outputs. While the model successfully created an interactive 3D game where users can dress Barbie and visit a fitting room, it struggled significantly with rendering accurate hands, facial features, and body proportions, leading the creator to conclude that AGI has not yet arrived.
Warp’s Productivity dashboard is an eng manager’s dream
Warp's new productivity dashboard centralizes engineering team visibility for managers and CTOs, enabling real-time monitoring of development velocity, code quality improvements, and cost efficiency across distributed teams rather than relying on individual local setups.
She uses Claude Code to find a house with the right vibe
Hillary shares an AI workflow she uses for house hunting that leverages Claude Code to find listings and filter them against her subjective "vibe-specific criteria," which she describes as evaluating whether a space would feel good to wake up in and spend time at.
She hasn’t sent a birthday text herself in 6 months
A speaker shares how they've delegated birthday acknowledgments to AI for the past 6 months, having their AI assistant research friends' children's ages and purchase appropriate gifts on Amazon. They frame this as automating 'emotional labor' through AI technology.
JJ’s Claude Code skill sets up new projects for him
JJ discusses his favorite meta workflow for setting up new projects using Claude, which automates the creation of repos, folders, Claude MD files, and project structure. This skill significantly reduces setup time by leveraging trained knowledge on repository building practices.