On GPT 5.5: Most ChatGPT users don't have problems complex enough to justify its cost
A speaker demonstrates GPT 5.5 building an educational app for teaching second-grade subtraction, questioning whether 17+ minutes of AI processing time is necessary for such simple tasks. They argue that most users don't have problems complex enough to justify this level of computational power and suggest the interface may not be optimal for non-technical users.
Summary
The speaker shares their experience using GPT 5.5 to create an educational application for teaching advanced subtraction concepts to their second grader. The AI system spent 17 minutes and 27 seconds thinking through the problem before generating a complete app with lessons, word problems, and read-aloud features. While acknowledging the app's functionality, the speaker questions the efficiency and necessity of such extensive processing time for relatively simple tasks. They express skepticism about whether this represents the right form factor for non-technical software engineers to access AI capabilities. The speaker concludes by noting their preference for using GPT models to solve complex technical problems rather than front-end development tasks, implying that the current implementation may be overpowered for typical user needs and questioning the value proposition of such computationally intensive AI for everyday applications.
Key Insights
- The speaker questions whether 17 minutes of hyper-intelligent AI thinking is necessary to build a simple educational app for second-grade subtraction
- The speaker expresses uncertainty about whether the current form factor is appropriate for non-technical software engineers to access AI capabilities
- The speaker states they don't typically use GPT models for front-end development work
- The speaker prefers to use AI models specifically for solving their hardest technical problems rather than simple application building
- The speaker demonstrates that GPT 5.5 can successfully create educational apps with multiple features including lessons, word problems, and read-aloud functionality
Topics
Transcript
[0:00] I'm teaching my second grader two-digit and three-digit subtraction. One of the ways that I've been able to teach him is build these little apps. And so I asked it to build an app for me to teach my second grader more advanced subtraction concepts. First out the gate, it's a thinker. So you can see here it thought for 17 minutes 27 seconds about this. It planned a app for advanced subtraction, built the code, all this kind of stuff. Now, here's my question. Do we need 17 minutes of hyper-intelligent thinking to build this app? Probably not. What are [0:30] we going to do with all this intelligence? Is this the right form factor for a non-technical…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
GPT-5.6's video editing via Codex is genuinely one of my favorite new workflows
A product manager describes using GPT-5.6 with Codex to automate video editing for social media clips. By simply dragging a file and providing natural language instructions, the AI generated five polished, fast-paced hype videos from a long conference talk recording, dramatically reducing the time-intensive manual clipping process.
Theoretically Intelligent vs. Practically Effective: Why GPT-5.6 Sol Beats Fable for Product Work
An executive contrasts two AI models (Fable and Soul/GPT-5.6 Sol), arguing that Soul is superior for product work because it prioritizes practical effectiveness over theoretical intelligence. The speaker values the ability to ship products to customers and understand end-user goals over theoretical sophistication.
Build a harness when the same workflow needs the same setup and the same outcomes, every time
Building a harness for repetitive workflows allows you to be more prescriptive about job execution, resulting in greater efficiency, consistency, and better outcomes. Rather than explaining requirements to an AI agent each time, a harness lets you use a simpler interface like pasting a link while the agent already understands the intended task.
GPT 5.6-Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark
Claire Vo compares OpenAI's new GPT 5.6 models (Soul, Terra, Luna) against Claude's Fable using her custom "How I AI" benchmark, finding that GPT 5.6 Soul excels at practical product work, prototyping, and natural communication, while Fable is theoretically intelligent but pedantic and difficult to collaborate with.
Context offloading is an underrated AI use case
The speaker highlights context offloading as an underrated AI use case, where AI serves as a safety net for routine cognitive tasks like email management and personal finances. Rather than adding new capabilities, AI reduces anxiety about missing important information or making mistakes by handling monitoring tasks, thereby freeing up mental bandwidth.