Opus 5: too timid?
The speaker criticizes Opus 5 for being excessively cautious and indecisive, constantly asking for human permission or approval rather than taking action. Even when faced with simple tasks like fixing a one-line merge conflict, Opus 5 deferred to the original branch owner's preferences, requiring the speaker to repeatedly push it to make autonomous decisions.
Summary
The speaker describes a frustrating pattern of interaction with Opus 5, an AI system that exhibits excessive timidity and hesitancy in problem-solving. Rather than confidently tackling tasks, Opus 5 repeatedly questions whether it should take action, asking variations of 'do you think I should do this, or should you, or should we ask someone else?' This behavior was consistent throughout their work together. A specific example illustrates this problem: when the speaker asked Opus 5 to resolve a trivial one-line merge conflict, instead of simply fixing it, Opus 5 raised concerns that the branch belonged to someone else and that resolving the conflict without that person's knowledge could disrupt their local work. The speaker found this response frustrating and patronizing, viewing it as an excuse rather than a legitimate concern. Throughout the entire interaction period, the speaker experienced this same pattern repeatedly—Opus 5 would identify solutions but lack confidence to implement them, often requesting human verification or approval. The speaker's consistent response was to override these hesitations and directly instruct Opus 5 to proceed with the task. This reveals a fundamental issue with Opus 5's design or behavior: it prioritizes caution and deference to a degree that undermines its usefulness and efficiency as a tool.
Key Insights
- Opus 5 consistently avoided making decisions by asking whether the user wanted to handle tasks instead, requiring constant user pushback to take action
- When asked to fix a simple one-line merge conflict, Opus 5 refused based on concerns that the branch belonged to someone else and resolving it could disrupt their local work
- Opus 5 repeatedly requested human verification and approval for actions it could autonomously perform
- The speaker's core complaint is that Opus 5 was excessively deferential—identifying correct solutions but lacking confidence to execute them without permission
- The pattern of excessive caution was not occasional but systematic, requiring the user to repeatedly override Opus 5's hesitations with direct commands
Topics
Transcript
[0:00] Let me just give an example of its timidity, where it was like, I think this is the answer, but do you think I should do it, or do you want to do it, or should we ask someone else to do it? It was like, every time I just kept saying like, why don't you solve this? Why don't you do this? And this is a really good example. I pulled a branch, and I was like, there is truly like a one-line merge conflict. Like, can you fix this merge conflict? And it was like, oh, but that's someone else's branch. Like, that's not my branch. I don't want to do that without him [0:30] knowing. It's…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
GPT-5.6's video editing via Codex is genuinely one of my favorite new workflows
A product manager describes using GPT-5.6 with Codex to automate video editing for social media clips. By simply dragging a file and providing natural language instructions, the AI generated five polished, fast-paced hype videos from a long conference talk recording, dramatically reducing the time-intensive manual clipping process.
Theoretically Intelligent vs. Practically Effective: Why GPT-5.6 Sol Beats Fable for Product Work
An executive contrasts two AI models (Fable and Soul/GPT-5.6 Sol), arguing that Soul is superior for product work because it prioritizes practical effectiveness over theoretical intelligence. The speaker values the ability to ship products to customers and understand end-user goals over theoretical sophistication.
Build a harness when the same workflow needs the same setup and the same outcomes, every time
Building a harness for repetitive workflows allows you to be more prescriptive about job execution, resulting in greater efficiency, consistency, and better outcomes. Rather than explaining requirements to an AI agent each time, a harness lets you use a simpler interface like pasting a link while the agent already understands the intended task.
GPT 5.6-Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark
Claire Vo compares OpenAI's new GPT 5.6 models (Soul, Terra, Luna) against Claude's Fable using her custom "How I AI" benchmark, finding that GPT 5.6 Soul excels at practical product work, prototyping, and natural communication, while Fable is theoretically intelligent but pedantic and difficult to collaborate with.
Context offloading is an underrated AI use case
The speaker highlights context offloading as an underrated AI use case, where AI serves as a safety net for routine cognitive tasks like email management and personal finances. Rather than adding new capabilities, AI reduces anxiety about missing important information or making mistakes by handling monitoring tasks, thereby freeing up mental bandwidth.