OpinionTechnical

LIVE: How I AI Opus 5.5 vs. GPT-6 Sol

How I AI

A live streaming review comparing three AI models (Claude Opus 5.5, GPT-6 Soul, and GPT-6 Astra) across multiple tasks including coding, design, SVG creation, and agent personality. The host conducted blind testing with custom benchmarks and found that while Opus 5.5 excels at agent tasks, the GPT models (particularly Astra and Soul) perform better for creative work and user experience.

Summary

The host conducts a real-time, live-streamed comparison of three AI models: Claude Opus 5.5, GPT-6 Soul, and GPT-6 Astra. The stream encounters technical difficulties with audio/video but proceeds as a DIY production. The host explains that these are not frontier-level models but rather cost-optimized versions with improved speed and token efficiency through caching optimization.

The evaluation framework has been significantly expanded beyond previous benchmarks. Rather than relying solely on standard benchmarks, the host tests models across diverse real-world tasks: PRD generation, personal productivity (email sorting and response), frontend design (multiple UI prototypes), backend code quality, agent personality and conversation style, long-form research tasks, SVG and character illustration creation, and video editing. The host emphasizes the subjective, 'vibrational' nature of assessment, acknowledging personal preference as a core evaluation metric alongside LLM-as-judge scoring.

Key findings from the blind testing reveal: (1) Opus 5.5 represents a significant improvement in conversational quality compared to Opus 5, losing the annoying verbosity that previously drove the host away from Claude models; (2) GPT-6 Soul emerges as the best value proposition—cheap and fast while maintaining quality; (3) Astra and Soul excel at creative tasks like SVG and character illustration, while Claude's SVG outputs tend toward repetitive design patterns (circles in corners, dense layouts); (4) Claude Opus 5.5 performs strongest in agent voice, long-form tasks, and backend code, while GPT models dominate in readability and design aesthetics; (5) All models struggle significantly with video editing and 3D rendering (demonstrated humorously through a Barbie 3D model test).

The host's personal preference clearly favors the GPT models, particularly for daily use, despite acknowledging Opus 5.5's genuine improvements. The analysis reveals a tension between benchmark performance and subjective user experience—Claude may score higher on certain technical metrics but produces outputs the host finds less enjoyable to interact with. The host notes that personality and risk management philosophy of AI companies seep into model behavior beyond just safety-critical tasks.

Key Insights

  • Opus 5.5 resolves the primary usability complaint about Opus 5 by reducing verbosity and becoming less annoying in conversation, returning Claude to the host's daily workflow after 60-90 days of preferring GPT models
  • GPT-6 Soul offers the best value proposition as a daily driver model—it is significantly cheaper than Opus 5.5 while maintaining quality performance and delivering faster response times
  • All tested frontier and non-frontier models struggle severely with creative video editing and 3D rendering tasks, suggesting these remain unsolved problems across different model architectures
  • Claude models tend to generate dense, complex frontend designs with repetitive pattern elements (circles, heavy text weight), while GPT models produce cleaner, more readable UI designs with better use of color and whitespace
  • The host's subjective preferences (which models feel enjoyable to use) frequently diverge from LLM-as-judge technical scoring, with the judge rating Fable highly while the host dislikes it, suggesting personal ergonomics matter more than raw performance metrics

Topics

AI model comparison and benchmarkingClaude Opus 5.5 vs GPT-6 modelsCost optimization and token efficiencySubjective vs objective model evaluationFrontend and UI design generationBackend code qualitySVG and creative illustrationAgent personality and conversation styleModel safety and risk management approachesReal-world AI application testing

Transcript

[0:00] and no one is helping me here . So we really work alone. Me, OpenAI, Anthropic , Claude, Codex, and the YouTube gods – this is , um, what we do. So I'm really excited to see how it goes and we're going to go live. I'm going to share my screen and we'll just see how it goes. Good. So , three models, one week. This is my life. Honestly, I got up this morning to [0:30] record the episodes and I thought, "I can't do three episodes today." So we'll just try it live. Um, Claude Opus 5.5, GPT-6 Soul and GPT-6 Luna. I was able to do some early testing of all of these models, so…

Full transcript available for MurmurCast members

Sign Up to Access

More from How I AI

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.