Codex Shifted Categories Entirely (and Nobody Noticed)
OpenAI transformed Codex from a coding tool into a comprehensive desktop agent that can control any Mac application through visual interface interaction, significantly outperforming Claude's computer use capabilities. This represents a strategic shift toward computer work rather than just knowledge work, enabled by acquiring a specialized team with deep macOS expertise.
Summary
OpenAI completely revamped Codex on April 16th, transforming it from a simple command-line coding tool into a sophisticated desktop agent capable of operating any Mac application through visual interface control. The transformation occurred in stages throughout 2025, with the Mac desktop app launching in February, Windows support in March, and the major computer use capabilities arriving in April. Codex can now see screens, click and type like a human, run multiple agents in parallel in the background, generate images, browse the web, and maintain memory across sessions.
The speaker extensively compares Codex to Claude's computer use capabilities, finding Codex significantly faster and more reliable. While Claude's cursor hesitates and sometimes requires retries, Codex operates at near-human speed and rarely fumbles tasks. This performance advantage stems from both GPT 5.4's native computer use capabilities and what OpenAI calls 'deep OS level wizardry' - background agents that don't hijack the user's cursor or steal focus.
The analysis reveals fundamental strategic differences between OpenAI and Anthropic. Anthropic focuses on knowledge work through structured interfaces, MCP servers, and explicit permissions, requiring ecosystem cooperation. OpenAI pursues broader 'computer work' through direct graphical interface control, eliminating the need for vendor cooperation or API development. This approach means any software with a graphical interface becomes automatable, including legacy enterprise software and internal tools.
Codex's computer use capabilities originated from OpenAI's October 2025 acquisition of Software Applications Incorporated, a 12-person team behind an unreleased Mac AI interface called Sky. This team previously built Workflow (acquired by Apple and turned into Shortcuts) and includes former Apple engineers with deep macOS expertise. The speaker emphasizes that both labs are acquiring specialized teams rather than just intellectual property, as model capabilities converge while human expertise remains scarce.
Looking forward, both companies aim for persistent, ambient, event-driven agents. OpenAI's Chronicle feature captures screen activity to improve computer use over time, while Anthropic's leaked Conway system represents an always-on agent environment with webhook triggers and browser control. The speaker predicts OpenAI's approach is more likely to succeed because it doesn't require ecosystem cooperation, though acknowledges enterprise software could move faster than expected in Claude's favor.
Key Insights
- Greg Brockman stated that models have shifted from being the product to being part of the product, with the focus now on building the 'body' around the AI brain
- GPT 5.4 benchmarks in the mid-70s on OS World, placing it above the human baseline for graphical user interface control
- Sam Altman admitted that OpenAI started the year behind Anthropic on real-world coding data and only appreciated the gap in hindsight
- OpenAI acquired the entire 12-person team from Software Applications Incorporated who built Workflow (which became Apple Shortcuts) and previously had deep Apple OS experience
- OpenAI leadership cut popular products like Sora and drug discovery efforts because they didn't align with their three strategic vectors: agentic platform, computer work, and personal AGI
Topics
Transcript
[0:00] OpenAI revamped Codeex completely and I am blown away by how useful that new app is. On April 16th, OpenAI turned Codeex into a desktop agent that operates every single app on your Mac. Clicking, typing, running in the background while you work. It's faster and more reliable than Claw's version of computer use, and it's by a much bigger margin than I expected. That matters because most enterprise software doesn't have modern APIs, and Codeex doesn't need them. I'm going to walk you through what Codeex is now, how it got this good, what OpenAI is really building, and where both labs [0:31] are going, and of course, what you should do about it. The reason to pay…
Full transcript available for MurmurCast members
Sign Up to AccessMore from AI News & Strategy Daily | Nate B Jones
The AI skill nobody talks about (and it isn't prompting) #AI #prompting #productivity #tech
The key differentiator in AI productivity isn't prompting skills but the ability to write structured specifications that enable AI to function as an autonomous agent. A person with advanced specification skills can produce 10x more output than someone using basic prompting by investing upfront time in detailed requirements and then letting the AI work independently.
1.6M agents registered for OpenClaw and did NOTHING.
The speaker explains how to determine whether a task requires a single agent, multiple agents, a chat interface, or no AI at all by using four key estimation criteria. He addresses the failure of 1.6 million OpenClaw agents that were registered but unused, arguing the problem is matching tasks to appropriate solutions rather than a lack of tools.
The one question that tells you if your role is safe #AI #careers #AIjobs #jobs #tech
The speaker presents a critical question for evaluating job security in the age of AI: would your role still exist if the company were significantly smaller? If the answer is no, your value is tied to coordination rather than direct value creation, making your position vulnerable in leaner organizations. The solution is to migrate toward work that directly generates revenue and drives business direction while adopting engineering principles of precision, testability, and falsifiability.
When everyone can code, this is what's scarce #AI #careers #AIjobs #coding #tech
As AI coding capabilities become widespread, the critical skill shifts from writing code to translating business needs into precise specifications and validating whether solutions actually solve customer problems. The person who can bridge vague requirements and technical implementation while exercising judgment becomes the organization's center of gravity.
20 AI Agents Rebuilt My Wife's Website For $8. I Never Typed a Word.
A developer demonstrates how a multi-agent AI system rebuilt his wife's website in 1.5 hours for $8 by orchestrating cheaper models under a premium supervisor, catching four major failures (hallucinations, accessibility shortcuts, design bugs, and checker errors) without human intervention—achieving superior results compared to six days of single-agent work.