LIVE: How I AI Opus 5.5 vs. GPT-6 Sol
A live streaming review comparing three AI models (Claude Opus 5.5, GPT-6 Soul, and GPT-6 Astra) across multiple tasks including coding, design, SVG creation, and agent personality. The host conducted blind testing with custom benchmarks and found that while Opus 5.5 excels at agent tasks, the GPT models (particularly Astra and Soul) perform better for creative work and user experience.
Summary
The host conducts a real-time, live-streamed comparison of three AI models: Claude Opus 5.5, GPT-6 Soul, and GPT-6 Astra. The stream encounters technical difficulties with audio/video but proceeds as a DIY production. The host explains that these are not frontier-level models but rather cost-optimized versions with improved speed and token efficiency through caching optimization.
The evaluation framework has been significantly expanded beyond previous benchmarks. Rather than relying solely on standard benchmarks, the host tests models across diverse real-world tasks: PRD generation, personal productivity (email sorting and response), frontend design (multiple UI prototypes), backend code quality, agent personality and conversation style, long-form research tasks, SVG and character illustration creation, and video editing. The host emphasizes the subjective, 'vibrational' nature of assessment, acknowledging personal preference as a core evaluation metric alongside LLM-as-judge scoring.
Key findings from the blind testing reveal: (1) Opus 5.5 represents a significant improvement in conversational quality compared to Opus 5, losing the annoying verbosity that previously drove the host away from Claude models; (2) GPT-6 Soul emerges as the best value proposition—cheap and fast while maintaining quality; (3) Astra and Soul excel at creative tasks like SVG and character illustration, while Claude's SVG outputs tend toward repetitive design patterns (circles in corners, dense layouts); (4) Claude Opus 5.5 performs strongest in agent voice, long-form tasks, and backend code, while GPT models dominate in readability and design aesthetics; (5) All models struggle significantly with video editing and 3D rendering (demonstrated humorously through a Barbie 3D model test).
The host's personal preference clearly favors the GPT models, particularly for daily use, despite acknowledging Opus 5.5's genuine improvements. The analysis reveals a tension between benchmark performance and subjective user experience—Claude may score higher on certain technical metrics but produces outputs the host finds less enjoyable to interact with. The host notes that personality and risk management philosophy of AI companies seep into model behavior beyond just safety-critical tasks.
Key Insights
- Opus 5.5 resolves the primary usability complaint about Opus 5 by reducing verbosity and becoming less annoying in conversation, returning Claude to the host's daily workflow after 60-90 days of preferring GPT models
- GPT-6 Soul offers the best value proposition as a daily driver model—it is significantly cheaper than Opus 5.5 while maintaining quality performance and delivering faster response times
- All tested frontier and non-frontier models struggle severely with creative video editing and 3D rendering tasks, suggesting these remain unsolved problems across different model architectures
- Claude models tend to generate dense, complex frontend designs with repetitive pattern elements (circles, heavy text weight), while GPT models produce cleaner, more readable UI designs with better use of color and whitespace
- The host's subjective preferences (which models feel enjoyable to use) frequently diverge from LLM-as-judge technical scoring, with the judge rating Fable highly while the host dislikes it, suggesting personal ergonomics matter more than raw performance metrics
Topics
Transcript
[0:00] and no one is helping me here . So we really work alone. Me, OpenAI, Anthropic , Claude, Codex, and the YouTube gods – this is , um, what we do. So I'm really excited to see how it goes and we're going to go live. I'm going to share my screen and we'll just see how it goes. Good. So , three models, one week. This is my life. Honestly, I got up this morning to [0:30] record the episodes and I thought, "I can't do three episodes today." So we'll just try it live. Um, Claude Opus 5.5, GPT-6 Soul and GPT-6 Luna. I was able to do some early testing of all of these models, so…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
Warp agents open PRs to fix the factory itself
Programming agents can autonomously improve factory systems by analyzing failed launches and proposing specific updates to agent definitions. A self-improvement loop enables observer agents to detect failures and generate evidence-based modifications that prevent recurring issues, such as changing specific steps in factory agent procedures.
Humans are still the bottleneck in Warp’s AI factory
Warp discusses how human code review has become the main bottleneck in their AI-assisted software development process, with a 3.5-hour delay from PR to first human review compared to 35 minutes from launch to PR. They're evolving their workflow to reduce human dependency by allowing requesters to review agent-generated code themselves, and plan to eventually skip review entirely for low-risk tasks by treating code review as a risk management exercise.
I Quit Claude Because It Was Annoying
The speaker explains why they stopped using Claude, citing frustrations with its tendency to produce nonsensical output and communicate in an unnatural, non-human manner. They mention considering a switch to Opus 5.5 but remain uncertain about fully migrating their work tasks.
Claude Is Not a Party Boy
A humorous character description of Claude as someone with traditional values who prioritizes work over social indulgence. The transcript portrays Claude as principled, occasionally frustrating, and willing to push back on tasks he finds objectionable.
Claude Is Back. I Still Reach for Codex.
The speaker explains their preference for using Codex over Claude, citing superior tooling, a better desktop application, and specific strengths in front-end design and SVG work. Despite Claude's return, they continue to reach for Codex for their development needs.