Claude sucks. And here is why.
This video explores why Claude and other AI models produce inconsistent results, examining how model selection (Haiku vs Sonnet vs Opus vs Claude 3.5 Fable), effort levels, prompt clarity, and system context dramatically affect output quality. The speaker demonstrates these differences through a Venn diagram creation task and advocates for using organized folder systems with documented SOPs and code-based solutions to ensure consistent AI performance.
Summary
The speaker addresses a recurring community complaint: AI assistants like Claude forgetting context, going off-rails, and behaving inconsistently. The core thesis is that inconsistency stems from multiple overlapping factors rather than a single problem.
First, the speaker introduces their PKM (Personal Knowledge Management) system called 'my PKA,' where Claude is configured as 'Larry,' an orchestrator agent that delegates tasks to specialized team members based on documented SOPs, guidelines, and work streams. This folder-based approach helps maintain consistency by providing persistent context.
The video's main contribution is demonstrating how different variables affect AI output quality: (1) Which AI platform you use (Claude, Gemini, etc.), (2) Which model within that platform (Haiku, Sonnet, Opus, Fable), (3) What effort/compute level is selected, and (4) How detailed and clear the instructions are.
Through a practical test—asking different models to create a four-circle Venn diagram with straight lines forming squares—the speaker shows dramatic quality differences. Haiku fails completely at the task. Sonnet creates a basic working diagram. Opus adds unnecessary features like sliders. Fable elegantly solves the problem by rotating the diagram 90 degrees for optimal placement without being asked. This demonstrates that Fable understood the intent better than more-capable-appearing models.
The speaker emphasizes that consistency is nearly impossible with pure prompting because: (1) The same prompt run twice on the same model produces different outputs due to LLM stochasticity, (2) This means validation requires running tests three times to confirm consistency, and (3) Most YouTube comparisons only run prompts once, making their conclusions unreliable.
To address inconsistency, the speaker recommends: using documented SOPs and guidelines within organized folders, converting critical workflows into code (which has deterministic outputs), creating specialized agents with documented expertise, using closed-session commands to create documentation and lessons learned, and asking AI to identify why it's deviating from its instructions.
The speaker notes that folder context only works in specific applications: Claude Desktop with Codex or Claude Code can write back to folders, but browser-based Claude cannot. They demonstrate that even with folder context, the folder doesn't automatically improve outputs unless it contains relevant instructions for that specific task.
Key Insights
- The same prompt run twice through the same model with identical settings produces completely different outputs due to how LLMs interpret and predict language differently each time, making single-test comparisons unreliable—proper validation requires running the same prompt three times to confirm consistency.
- Claude 3.5 Fable demonstrated superior understanding by rotating the Venn diagram 90 degrees without being asked, recognizing that this was the most efficient spatial arrangement, whereas Opus added unnecessary features like interactive sliders that weren't requested.
- Haiku model cannot even create a basic overlapping Venn diagram with squares and circles out of the box, and maxing out its effort level (extended thinking X high) still fails to produce the correct output, indicating fundamental capability limitations below a certain model tier.
- Folder context in Claude only works in specific applications (Claude Desktop with Codex or Claude Code that can write back to folders), not in the browser version—browser uploads are read-only and don't allow the AI to actually update files, defeating the purpose of folder-based systems.
- Converting workflows from pure prompting into code guarantees consistency because code produces deterministic mathematical outputs, whereas prompting always involves LLM prediction that varies—for critical outcomes, teams should identify workflow steps that can be automated as code rather than relying solely on prompts.
Topics
Transcript
[0:00] If you're using Claude and it keeps forgetting or not doing what you want, that is going off rail or behaving differently than it did yesterday. This video is for you because this is a recurring topic that comes up in our community over and over. And that's why I'm making this video to maybe give you some more perspective, why these things happen. So if you can relate to this sentence, AI intelligent and the most stupid assistant in the world, then this video is for you, just for context, every week inside my iCore, we launch a poll here. [0:35] In this case, it was, AI workflow most often break down and slow you down? And there…
Full transcript available for MurmurCast members
Sign Up to AccessMore from ICOR with Tom | AI Productivity
My AI team took 8 hours. Claude Sonnet 5.5 took 15 minutes.
A creator comparing an 8-hour AI team project with a 15-minute Claude Sonnet 5.5 solution reveals that excessive guardrails and restrictions paradoxically hindered the multi-agent system's performance. The speaker demonstrates how reducing constraints and context allows AI to produce more creative and functional results, though both approaches have significant limitations for production use.
Your knowledge should outlive every AI tool you use.
The speaker clarifies the distinction between the ICOR methodology (a tool-agnostic productivity framework) and its implementations like the ICOR for Life scaffold and myPKA AI team, emphasizing that users can adopt the core principles in any tool they prefer. They announce plans to restructure their membership platform by separating these components to reduce confusion and help users understand they're not locked into specific tools or systems.
Claude alone took 3 minutes. With Jev, 21 seconds.
A demonstration comparing AI slide deck creation using Claude alone versus Claude integrated with Jev, showing that the hybrid approach completed the task in 21 seconds for $8 with proper formatting, while Claude-only took 3+ minutes and cost $49 with poor formatting.
Why I stick to Claude for work (most of the time)
The creator demonstrates why Claude is their preferred AI for work by comparing it with ChatGPT across multiple tests, showing that Claude better adheres to custom instructions and agentic workflows defined in local folder structures. The key advantage lies in using organized, LLM-agnostic folder systems with agents.md files rather than relying on auto-memory features.
You are using the wrong Claude for work! (Here is proof)
A video demonstrating that Claude Code is significantly more powerful than Claude Cowork for professional knowledge workers, showing how Claude Code better understands folder context, enforces safety rules, orchestrates sub-agents with different models, and creates persistent file outputs rather than temporary artifacts.