Claude Opus 4.7 Is Crazy Good At Coding
Claude Opus 4.7 is Anthropic's latest model with significant improvements in agentic coding, instruction following, and multimodal support. It scores 64.3% on coding benchmarks, placing it between Opus 4.6 (53.4%) and Mythos preview (77.8%). The presenter recommends it as the go-to model for coding tasks in tools like Cursor or Claude Code.
Summary
The video covers the release of Claude Opus 4.7, Anthropic's newest model positioned as a top-tier option for coders. The presenter highlights that the most notable improvement is in agentic coding performance, citing benchmark scores to contextualize the leap: Opus 4.6 scored 53.4%, Mythos preview scored 77.8%, and the new Opus 4.7 lands in the middle at 64.3% — a meaningful step forward from its predecessor.
Beyond coding benchmarks, Opus 4.7 brings improvements in instruction following, multimodal support (better understanding of images), and memory handling. The presenter notes that for everyday Claude users, the most noticeable change will likely be in how well the model follows instructions. Older models reportedly required more careful prompt engineering and precise phrasing, whereas Opus 4.7 is described as more capable of handling instructions without that extra effort.
The presenter concludes by stating a personal preference to use Opus 4.7 going forward when writing code in tools like Cursor or Claude Code, citing the benchmark results as clear evidence of its superiority for coding use cases.
Key Insights
- The presenter argues that the biggest leap in Opus 4.7 over its predecessor is specifically in agentic coding, not general capability, as evidenced by the jump from 53.4% to 64.3% on coding benchmarks.
- The presenter notes that Opus 4.7 scores 64.3% on the agentic coding benchmark, placing it between Opus 4.6 at 53.4% and Mythos preview at 77.8%, framing it as a meaningful but not top-of-class performer.
- The presenter claims that older Claude models required more deliberate prompt engineering and precise phrasing, implying that Opus 4.7 reduces that burden for users.
- The presenter highlights improved multimodal support as a key feature of Opus 4.7, specifically noting better understanding of images alongside text.
- The presenter states a personal intention to use Opus 4.7 exclusively for coding tasks in tools like Cursor or Claude Code, citing benchmarks as the deciding factor.
Topics
Transcript
[0:00] If you are a coder, you now have a new best top-of-the-line model that you can be using Claude Opus 4.7, the brand new model out of Anthropic. Where the biggest leap seemed to be is in agentic coding. Opus 4.6 53.4%, Mythos preview 77.8%, a pretty huge leap. And this version they gave us this week, Opus 4.7, it kind of split right down the middle at 64.3%. It is substantially better at following [0:31] instructions. It's got improved multimodal support, so better at understanding images and things like that. And it's better with the memory here. If you're just like an everyday Claude user, you might notice the difference in instruction following. The way I sort of…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Matt Wolfe
#ad Why AI Intelligence Is Overrated
The transcript argues that raw AI intelligence is insufficient without operational guardrails and enterprise context. ServiceNow positions itself as an 'AI control tower' that embeds AI within enterprise systems with compliance, approval processes, and institutional knowledge—similar to how new engineers need oversight before accessing production systems.
AI News: GPT-5.6 and the new Super App are a Massive Leap!
OpenAI released GPT-5.6, a major leap forward in AI capabilities, alongside the new unified ChatGPT work app integrating code, browsing, and agent features. The week also saw significant model releases from xAI (Grock 4.5), Meta (Llama Spark 1.1), and research from Anthropic on AI reasoning patterns, establishing a new competitive landscape in AI development.
I Built A Monetizable Business With AI
The creator built a monetizable finance dashboard in one day using Hyper Agent from Airtable, which deploys a team of AI agents that continuously monitor news, track funding rounds and IPOs, analyze market sentiment, and generate daily briefings without manual intervention. The system demonstrates how specialized AI agents can work together autonomously to create a functional business product.
The ONLY AI Benchmark You Need!
A developer created "Buccy Bench," a humorous yet functional AI benchmark that tasks different language models with drawing Gary Busey as SVG code rather than images. The benchmark tracks model evolution over time while measuring performance metrics like cost, tokens, and execution speed.
GLM-5.2 - The Open Model That's As Good As Opus!
A comprehensive review of GLM-5.2, an open-weight Chinese AI model with a 1 million token context window, demonstrating its capabilities for coding, document analysis, and agentic workflows at significantly lower costs than frontier models like Claude Opus and GPT-4.5. The speaker tests various use cases including website building, Chrome extensions, game development, and data organization, concluding it's valuable for long, code-heavy, token-expensive tasks despite not universally outperforming closed-source alternatives.