AI News: The Scariest AI Model Ever!
This weekly AI news roundup covers Claude Mythos, an unreleased AI model from Anthropic that's so powerful at finding software vulnerabilities they won't release it publicly. Instead, they're giving limited access to major tech companies through Project Glass Wing to help patch vulnerabilities before similar models become widely available.
Summary
The video begins with extensive coverage of Claude Mythos and Project Glass Wing, which has dominated AI discussions this week. Mythos is Anthropic's most powerful unreleased AI model that demonstrates unprecedented coding capabilities, particularly in finding and exploiting software vulnerabilities. The model scored 83.1% on cybersecurity vulnerability reproduction benchmarks, significantly outperforming previous models, and found vulnerabilities in major operating systems and browsers, including a 27-year-old vulnerability in OpenBSD and a 16-year-old vulnerability in FFmpeg. Rather than releasing it publicly due to security concerns, Anthropic launched Project Glass Wing, giving selective access to major tech companies like Apple, Microsoft, and Nvidia so their cybersecurity teams can identify and patch vulnerabilities before similar powerful models become widely available.
The video then covers two major language model releases: Meta's Muse Spark from their new Super Intelligence Labs, which represents their first significant model since restructuring their AI team, and GLM 5.1 from ZAI, an open-source model under MIT license that achieves near state-of-the-art coding performance. Meta's model performs well across benchmarks, ranking fourth overall, while being highly token-efficient, though it doesn't lead in any particular category. The GLM 5.1 model is particularly noteworthy for matching GPT-4's coding performance while being completely open-source and available for download.
Google introduced new features for Gemini, including interactive simulations and models similar to OpenAI and Anthropic's recent releases, plus a new Notebooks feature that organizes chats and files like projects in other platforms. The video also covers the US rollout of Seed Dance 2.0 video model through Runway and CapCut, HeyGen's new Avatar 5 model requiring only 15 seconds of recording, and various other updates including OpenAI's new $100/month Pro tier, Anthropic's managed agents feature, and multiple smaller tool releases and features across the AI landscape.
Key Insights
- Anthropic built Claude Mythos as a general-purpose model for coding that accidentally became extremely powerful at cybersecurity, finding vulnerabilities in major operating systems without specific training for cyber tasks
- The speaker argues there's a pattern of companies claiming their models are 'too dangerous to release' for marketing purposes, comparing current Mythos headlines to similar claims about GPT-2 in 2019
- Meta's new Super Intelligence Labs produced their first model Muse Spark, which ranks fourth overall in AI benchmarks and is highly token-efficient, marking their first major release since restructuring their AI team
- GLM 5.1 achieves near state-of-the-art coding performance comparable to GPT-4 and Claude while being completely open-source under MIT license, which the speaker finds mind-blowing
- Anthropic announced they will no longer allow Claude subscriptions to cover usage in third-party tools like OpenClaw starting April 4th, likely because these tools burn through tokens faster than subscription revenue can cover
Topics
Transcript
[0:05] I just got back from the Human X event out in San Francisco last night. It was an event totally focused on the people and the companies building in the AI space. I met a lot of amazing people, but while I was gone, a ton of AI news happened. And well, this is your weekly deep dive into everything that you need to know that happened in the world of AI from the past week. I'm not going to waste your time, so let's dive right in. Let's start with Claude Mythos and Project Glass Wing. This is the story that literally everybody in the AI space [0:35] is talking about. I've seen probably six videos now of…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Matt Wolfe
AI News: Dots, GPT-6.1 Sol, Sonnet 5.5, Gemini 4, and everything you need to know
A comprehensive review of major AI announcements from the week, including OpenAI's new Dots agent, GPT-6.1 Soul model, Anthropic's Sonnet 5.5, and Google's Gemini 4 Argon, with analysis of pricing, capabilities, and competitive positioning across different models.
The Hands Down Best Coding Model Right Now
The speaker reviews Opus 5.5, claiming it's currently the best state-of-the-art coding model at a reasonable price point. They demonstrate its capabilities by running a 20-hour test where the AI built a nearly complete recreation of the Mega Bonk game, including accurate character tiers, sound design, and UI elements.
This AI Model Does’t Use Words?!
Typesafe AI released Jev, a new AI model that outputs structured decisions (choices, scores, or booleans) instead of generated text like traditional LLMs. This approach makes Jev dramatically cheaper and faster, with input costs at 4 cents per million tokens and free output tokens.
AI News: Opus 5.5, GPT-6 Sol, Jev, Muse and More!
This week in AI saw major announcements from Meta Connect (Muse agent, new AR/VR glasses), three new language models (Claude Opus 5.5, GPT-6 Soul, Grok 4.7), significant buzz around Typesafe AI's Jev decision model, and new features from YouTube, Microsoft, Google, and Spotify. Claude Opus 5.5 emerged as the new state-of-the-art model while Jev introduced a fundamentally different approach to AI outputs focused on decision-making rather than text generation.
5 Ways To Use ChatGPT's New Features
The transcript demonstrates five practical applications of ChatGPT's new features, including computer use for software automation, advanced image editing, cloud browser integration for account access, voice-powered live agents, and persistent session agents for long-term user interaction tracking.