NEW DeepSeek V4 Flash Update!
DeepSeek V4 Flash has been officially released with significant performance improvements over its preview version, designed specifically for AI agents. The speaker demonstrates practical applications by building 50+ projects and integrating it into their agent operating system, while emphasizing it's not frontier-level but offers excellent speed and cost efficiency.
Summary
DeepSeek announced a major upgrade to V4 Flash, transitioning from preview to official API status. The benchmark improvements are substantial: AIME increased from 61.8 to 82.7, NL2 repo from 39.4 to 54.2, Cyber Gym from 38.7 to 76.7, and Deep SWE from 7.3 to 54.4. The model performs comparably to Claude Opus 4.8 but is positioned below frontier models like Fable 5. Critically, DeepSeek achieved these improvements through better training rather than increasing model size—V4 Flash maintains the same architecture and size as its preview version. The speaker tested the model extensively by building over 50 projects in a single day, including interactive games and 3D environments, demonstrating its capability for complex logic and UI design. The model is specifically optimized for agentic work, making it ideal for AI agent frameworks like Hermes. Integration into the speaker's agent operating system took minutes, allowing the model to slot into existing workflows alongside Claude and other models. V4 Flash operates with a 1 million token context window and is 12% more token-efficient than its predecessor. Pricing is significantly lower than competitors—60% cheaper than GPT-5.6 and Luna Max. The speaker clarifies that the upgrade is API-only; the deepseek.com web interface and V4 Pro remain unchanged. Open weights are expected to release soon, which would make V4 Flash the second-highest scoring open-weight model behind Llama 3. The speaker emphasizes building flexible AI systems where new models can be plugged in immediately upon release rather than chasing individual models.
Key Insights
- DeepSeek achieved significant performance upgrades to V4 Flash through improved training methodology rather than increasing model size or parameters, maintaining identical architecture to the preview version
- V4 Flash is specifically tuned for agentic work and agent frameworks like Hermes, not designed as a flagship coding model, representing a different optimization priority than frontier models
- The speaker built over 50 functional projects including 3D games in a few hours with V4 Flash, demonstrating practical capability for complex logic, UI design, and rendering despite not reaching frontier model quality
- V4 Flash is priced 60% lower than GPT-5.6 and comparable frontier models while achieving performance within one point of GLM 512 and comparable to Gemini 3.6 Flash
- The upgrade is API-only as of the release date; the deepseek.com web interface and V4 Pro API remain unchanged, with V4 Pro expected to release separately as a potential frontier-level model
Topics
Transcript
[0:00] A brand new version of DeepSeek, DeepSeek 4 Flash just went live and is built for AI agents. You can see the announcement just happened a few hours ago today. DeepSeek just dropped a major upgrade to V4 Flash and makes your agents seriously more powerful. So, DeepSeek say the new benchmark scores far surpass their previous top preview model from the small, fast, and cheap tier. I'll show you exactly what I've built with We actually built over 50 things with it already today. So, we've tested it relentlessly and that means your agents get to think across a [0:30] million tokens of context, they run longer coding loops, and finish more work uh before you touch your…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Julian Goldie SEO
Impeccable Makes Claude, Codex and Kimi 10X Better
Impeccable is a free, open-source design tool that improves AI-generated web designs by eliminating generic templates and providing design direction through product/design files and 23 specific commands. The tool works across Claude, Codex, and other models, detecting and fixing common AI design flaws like repetitive gradients, poor spacing, and accessibility issues before deployment.
4 FREE Repos to Cut Claude Code Tokens by 80%!
The video presents a token minimization playbook featuring four free open-source tools (RTK, Caveman, Ponytail, and Omni Root) designed to reduce Claude API token usage by up to 80%. These tools work together to filter unnecessary tool output, compress replies, reduce code bloat, and delegate grunt work to cheaper models, allowing users to get more efficiency from their AI agent systems.
This NEW Chinese AI is INSANE! (FREE + Open Source)
Kimi Chat 3 from Moonshot AI is a free, open-source Chinese AI model with 2.8 trillion parameters that can handle 1 million tokens in context and demonstrates strong performance across reasoning tasks, code generation, and real-time application building comparable to premium models.
NEW Google Search Console Update is INSANE!
Google has launched a new Search Console feature allowing users to connect social media profiles directly, enabling them to see which keywords their social posts rank for. This creates an opportunity to build a 'keyword empire' by taking proven keywords from Search Console and creating unique platform-native content across multiple social channels to rank multiple times for the same keyword.
Sakana Fugu Ultra 1 1 Review Masterclass
A review of Sakana Fugu Ultra 1.1, an AI model that uses a council of multiple expert models working together to generate higher-quality outputs than single models like Claude Fable 5. The model excels at complex tasks like game development but takes approximately 25 minutes per generation and has regional availability restrictions.