NEW DeepSeek V4 Flash Update!
DeepSeek V4 Flash has been officially released with significant performance improvements over its preview version, designed specifically for AI agents. The speaker demonstrates practical applications by building 50+ projects and integrating it into their agent operating system, while emphasizing it's not frontier-level but offers excellent speed and cost efficiency.
Summary
DeepSeek announced a major upgrade to V4 Flash, transitioning from preview to official API status. The benchmark improvements are substantial: AIME increased from 61.8 to 82.7, NL2 repo from 39.4 to 54.2, Cyber Gym from 38.7 to 76.7, and Deep SWE from 7.3 to 54.4. The model performs comparably to Claude Opus 4.8 but is positioned below frontier models like Fable 5. Critically, DeepSeek achieved these improvements through better training rather than increasing model size—V4 Flash maintains the same architecture and size as its preview version. The speaker tested the model extensively by building over 50 projects in a single day, including interactive games and 3D environments, demonstrating its capability for complex logic and UI design. The model is specifically optimized for agentic work, making it ideal for AI agent frameworks like Hermes. Integration into the speaker's agent operating system took minutes, allowing the model to slot into existing workflows alongside Claude and other models. V4 Flash operates with a 1 million token context window and is 12% more token-efficient than its predecessor. Pricing is significantly lower than competitors—60% cheaper than GPT-5.6 and Luna Max. The speaker clarifies that the upgrade is API-only; the deepseek.com web interface and V4 Pro remain unchanged. Open weights are expected to release soon, which would make V4 Flash the second-highest scoring open-weight model behind Llama 3. The speaker emphasizes building flexible AI systems where new models can be plugged in immediately upon release rather than chasing individual models.
Key Insights
- DeepSeek achieved significant performance upgrades to V4 Flash through improved training methodology rather than increasing model size or parameters, maintaining identical architecture to the preview version
- V4 Flash is specifically tuned for agentic work and agent frameworks like Hermes, not designed as a flagship coding model, representing a different optimization priority than frontier models
- The speaker built over 50 functional projects including 3D games in a few hours with V4 Flash, demonstrating practical capability for complex logic, UI design, and rendering despite not reaching frontier model quality
- V4 Flash is priced 60% lower than GPT-5.6 and comparable frontier models while achieving performance within one point of GLM 512 and comparable to Gemini 3.6 Flash
- The upgrade is API-only as of the release date; the deepseek.com web interface and V4 Pro API remain unchanged, with V4 Pro expected to release separately as a potential frontier-level model
Topics
Transcript
[0:00] A brand new version of DeepSeek, DeepSeek 4 Flash just went live and is built for AI agents. You can see the announcement just happened a few hours ago today. DeepSeek just dropped a major upgrade to V4 Flash and makes your agents seriously more powerful. So, DeepSeek say the new benchmark scores far surpass their previous top preview model from the small, fast, and cheap tier. I'll show you exactly what I've built with We actually built over 50 things with it already today. So, we've tested it relentlessly and that means your agents get to think across a [0:30] million tokens of context, they run longer coding loops, and finish more work uh before you touch your…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Julian Goldie SEO
How to Run DeepSeek V4 Flash for FREE!
A tutorial demonstrating how to use DeepSeek V4 Flash for free through Open Code and integration with agent operating systems like Hermes Agent. The speaker showcases building websites and apps using this free AI model and explains how it compares favorably to larger models despite being smaller.
Microsoft Fara1.5 27B NEW Browser Automation Model is WILD!
Microsoft released Phi-3.5, a family of three computer use models (4B, 9B, 27B) that automate browser tasks through vision-based clicking rather than HTML parsing. These open-weight models significantly outperform larger closed-source alternatives like OpenAI's Operator and Google's Gemini 2.0 on web automation benchmarks.
Claude Obsidian 2.0 is INSANE (FREE!)
Claude Obsidian 2.0 is presented as a free AI memory upgrade that allows users to upload files into a folder for permanent retention and linking. The system creates a knowledge graph that learns from business documents, provides sourced answers, and can be shared across teams.
Claude Agent OS is INSANE! 🤯
Julian presents a comprehensive Claude-based agent operating system that integrates multiple AI models, automated workflows, and a persistent memory system to automate daily tasks. The system runs 24/7 and uses free or existing subscriptions, combining tools like voice agents, content creation, competitor monitoring, and real-time news analysis into a single unified dashboard.
NEW ChatGPT Update is INSANE!
OpenAI released a major ChatGPT update featuring a Chrome extension called Side Chat and an improved desktop app that work together to streamline SEO research and content creation. The update allows users to analyze multiple browser tabs simultaneously, highlight text for quick answers, and convert research into finished work without constant tab switching.