NewsOpinion

Opus 5.5 Is Crazy Good and GPT-6-Sol Launched Too

Matt Wolfe

Anthropic released Claude Opus 5.5, a new state-of-the-art model that outperforms their previous best model (Fable 5.1) while being 40% cheaper and faster. OpenAI simultaneously launched GPT-6 Soul and Luna, which are cheaper and faster but less capable than their flagship Astra model, making Opus 5.5 the more impressive release of the day.

Summary

On the same day, two major AI model releases occurred. Anthropic announced Claude Opus 5.5 in the morning, claiming it performs at the level of Claude Fable 5.1 while costing 40% less to operate. The speaker reviews benchmark data showing Opus 5.5 achieving state-of-the-art results across multiple metrics including agentic coding (66.4%), frontier code (54.4%), and computer use tasks, surpassing both Fable 5.1 and GPT-6 Astra in most areas. Pricing dropped significantly: input tokens from $5 to $4 per million, and output tokens from $25 to $20 per million, making it substantially cheaper than Fable 5.1 ($10/$50). Artificial Analysis ranked Opus 5.5 as the objectively smartest model with a score of 58, compared to Fable 5.1 and GPT-6 Astra both at 53.

The speaker highlights impressive demonstrations of Opus 5.5's capabilities, showcasing animations, games, and interactive applications created by developers using the model. Examples include JavaScript animations, a Game Boy emulator, Blender-generated graphics, first-person shooter games, a Dark Souls game recreation, and a flight simulator—all created with Claude Opus 5.5. The speaker emphasizes that games and animations represent the best way to demonstrate model capabilities, noting the dramatic improvement from what these models could create a year prior.

OpenAI released GPT-6 Soul and Luna approximately two hours after Opus 5.5's announcement. GPT-6 Soul represents an incremental improvement with significantly reduced pricing (input tokens halved from $4 to $2, output tokens halved from $20 to $10), but it falls short of state-of-the-art performance compared to Astra. GPT-6 Luna is even cheaper but performs at an even lower tier. The speaker notes minimal benchmark data from OpenAI's announcement and limited demo circulation, suggesting the company anticipated weaker results compared to Opus 5.5. On Artificial Analysis, GPT-6 Soul scored 48, tied with Muse Spark, while GPT-6 Luna tied with GPT 5.6 Luna.

The speaker characterizes Opus 5.5 as the clear winner of the day, describing it as a rare case where a new model is simultaneously faster, cheaper, and more capable than its predecessor. OpenAI's release is framed as a marginal improvement—cheaper and faster than Astra, but not substantially better, making it primarily useful for cost-conscious developers rather than representing a meaningful capability advancement.

Key Insights

  • Claude Opus 5.5 achieves state-of-the-art performance on agentic coding benchmarks (66.4%) and beats both Fable 5.1 and GPT-6 Astra across most benchmark categories despite being positioned as a lower-tier model
  • Opus 5.5 costs significantly less than Fable 5.1 ($4 vs $10 per million input tokens, $20 vs $50 per million output tokens) while performing better on most tasks, making it superior on both capability and price dimensions
  • Developers have created remarkably sophisticated applications with Opus 5.5 including Game Boy emulators, Dark Souls game recreations, flight simulators, and complex JavaScript animations—capabilities the speaker notes were not possible with AI models a year ago
  • GPT-6 Soul uses significantly fewer output tokens per task (31,000) compared to Opus 5.5 (119,000), resulting in lower cost per task ($16 vs $5.98) despite using the same evaluation metrics
  • OpenAI's GPT-6 Soul release appears deliberately timed to maintain media presence following Anthropic's Opus 5.5 announcement, though it represents only a marginal improvement over Astra and lacks the early access demonstrations that accompanied Fable's release

Topics

Claude Opus 5.5 release and capabilitiesModel pricing and cost efficiency comparisonBenchmark performance analysisDeveloper demonstrations and use casesGPT-6 Soul and Luna releasesAI model capability evaluationGames and animations as capability demonstrations

Transcript

[0:00] Another day, another new best model in the world just came out. And of course, on the day where we get not one but two state-of-the-art models out of the big foundation labs, I'm out in Palo Alto for Meta Connect. So, doing a quick review of what just came out from my hotel room. So earlier this morning, we got a new model out of Anthropic, and it appears to be the new best of the best, better than Fable 5.1, but cheaper and faster. And this model is Claude Opus [0:35] 5.5 out of Enthropic. They claim on their website it now performs at the level of Claude 5.1 on most work and costs 40% less to…

Full transcript available for MurmurCast members

Sign Up to Access

More from Matt Wolfe

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.