OpinionTechnical

I reviewed Opus 5.5 and GPT-6 Sol live - and the results surprised me

How I AI

A live review comparing three newly released AI models—Opus 5.5, GPT-6 Soul, and GPT-6 Luna—where the reviewer conducts blind testing across multiple task categories and finds that while Opus 5.5 significantly improves on previous versions, OpenAI's models excel in creative tasks like SVG generation.

Summary

The reviewer conducted a comprehensive live evaluation of three new AI models released the same morning: Claude Opus 5.5 by Anthropic, GPT-6 Soul, and GPT-6 Luna by OpenAI. The reviewer had early access to Opus 5.5 and expanded their 'How I AI' benchmark to test models across diverse categories including PRDs, personal productivity, frontend coding, backend coding, agent personality, long-term research, computer usage, and creative tasks like SVG illustration and video editing.

Key findings about pricing and performance: Opus 5.5 costs nearly twice as much as GPT-6 Soul, though both are cheaper than their predecessors (Opus 5 and Fable 1). All three models showed improvements in cost efficiency, speed, and token optimization through cached inputs. Anthropic implemented Fable-level security restrictions in Opus 5.5 for cybersecurity and biology tasks.

Regarding user experience and communication style: The reviewer noted significant improvements in Opus 5.5's interaction quality compared to Opus 5, describing previous versions as 'annoying' and 'conservative.' Opus 5.5 produces clearer bulleted lists and normal behavior, though it remains somewhat verbose and provides less commentary than OpenAI models. GPT-6 Soul was found to be faster in actual latency and provided better balance in communication level.

Testing results showed distinct model strengths: For PRD writing, Soul performed best. For agent personality and long-term tasks, Claude Opus 5.5 excelled. For personal productivity tasks, Model B and Model E (which appeared to be different OpenAI variants) wrote the most readable messages. Surprisingly, OpenAI models significantly outperformed Claude on SVG character creation, producing cuter and more detailed illustrations. For frontend design, Claude models tended to over-complicate interfaces while OpenAI models provided cleaner, more usable designs. All models performed poorly on video editing tasks.

The reviewer's overall assessment: Opus 5.5 won 'the week' by receiving the highest average scores across individual tasks, particularly excelling in agent voice work and specific design tasks. However, the reviewer's personal preference ('my heart') went to OpenAI's Astra and Soul models, particularly due to Soul's cheaper pricing and superior performance on creative tasks. Notably, an LLM judge disagreed with the reviewer's assessments, preferring Fable and ranking Soul lower, highlighting differences between human and AI-based evaluation criteria.

Key Insights

  • Opus 5.5 represents a significant improvement in user experience and interaction quality compared to Opus 5, to the point where the reviewer stopped feeling irritated during everyday use and returned it to their daily workflow after months of preferring other models.
  • Anthropic's design choices to make Opus 5.5 less verbose by reducing commentary about its work creates an unintended user experience issue where people question whether the model is actually working, making it feel slower than it actually is.
  • OpenAI models (Astra and Soul) significantly outperform Claude models at SVG character creation and creative illustration tasks, contradicting the reviewer's initial prediction that Claude would excel at visual creative work.
  • Claude models tend to over-complicate frontend designs with excessive detail and richness when given complex requests, making prototypes difficult to understand, whereas OpenAI models maintain cleaner, more readable interfaces.
  • An LLM judge evaluating the reviewer's own test results disagreed substantially with their assessments, preferring Fable over Astra and rating Soul much lower, revealing fundamental differences in how AI judges versus humans evaluate model outputs.

Topics

AI model comparison and benchmarkingClaude Opus 5.5 improvements and limitationsGPT-6 Soul and Luna pricing and performanceBlind testing methodology and evaluation biasSVG generation and creative task performanceUser experience and communication style differencesFrontend and backend code generation capabilitiesModel pricing and cost efficiencyToken optimization and cachingSecurity restrictions in AI models

Transcript

[0:00] I was as prepared as possible for Anthropic to release Opus 5.5. I was even prepared for OpenAI to release another model. I got up early this morning because I had early access to Opus 5.5 to record an incredible review for you . And guess what? They both, they both came out this morning. So now, even though I already have a great Opus 5.5 review ready, [0:30] which you'll see later, I'm just going to do this live stream. And I've never done live broadcasts before. So I'm going to tell you about these new models. Three models were released today: Opus 5.5 by Anthropic, GPT-6 Soul, and GPT-6 Luna. There are price wars going on here,…

Full transcript available for MurmurCast members

Sign Up to Access

More from How I AI

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.