DeepSeek V4 Flash 0731 + Hermes Agent is INSANE!
DeepSeek V4 Flash, a newly released Chinese AI model optimized for agentic tasks, integrates with Hermes Agent to enable powerful automation workflows. The model offers significant performance improvements over its predecessor, cheaper pricing than frontier models, and a 1M token context window ideal for multi-step agent operations.
Summary
The video discusses DeepSeek V4 Flash (0731 build), a new model update that just became available through OpenRouter integration with Hermes Agent. The speaker clarifies that while initially announced for NewsPortal, the platform doesn't allow models trained on user data, so OpenRouter is the correct integration point. DeepSeek V4 Flash represents a major performance leap compared to the older Flash preview version, with superior agentic benchmarks. The model is cheaper than frontier alternatives like Claude Opus while maintaining comparable performance. The 0731 build is specifically a post-training upgrade focused on agent capability, featuring the same parameter size as before but with significantly better training for agent loops and tool calling. The speaker emphasizes that what previously limited agents wasn't Hermes itself, but the underlying model—factors like slow responses, poor tool-calling ability, memory loss over long jobs, and high frontier pricing all affected performance. DeepSeek V4 Flash addresses these through its 1M token context window, agentic-focused training, speed, and affordability. The speaker demonstrates three major use cases: creating AI avatar videos with full editing and voiceover, generating blog posts through multi-agent orchestration, and direct chat-based interactions. The content highlights the importance of building modular systems (Agent OS) rather than depending on specific models, as new models constantly emerge. The speaker promotes AI Profit Boardroom, a community offering the Agent OS, daily video tutorials, 24/7 support, and weekly calls for people building AI agent workflows.
Key Insights
- DeepSeek V4 Flash is specifically designed for agentic tasks with post-training focused on agent capability and tool-calling loops, making it fundamentally different from general-purpose frontier models despite being the same parameter size as the previous version
- The bottleneck for agent performance has historically been the underlying language model, not the Hermes agent framework itself—factors like model memory loss over 20+ steps, frontier pricing for each tool result, and lack of agentic training limited agent effectiveness
- DeepSeek V4 Flash's 1 million token context window is critical for multi-step agent jobs, allowing the model to remember everything throughout extended workflows without forgetting previous steps
- Building a modular Agent Operating System that can plug different models in and out is more valuable than optimizing for any single model, since new models constantly emerge and a flexible system ensures workflows never break
- DeepSeek V4 Flash response speed is significantly faster than frontier models like Claude through OpenRouter, delivering replies within seconds compared to 10-15 seconds, resulting in 2-3x more throughput in chat interactions
Topics
Transcript
[0:00] Deep Seek V4 Flash, the brand new update that just dropped yesterday, is now available on Hermes. I've tested I'm going to show you exactly how we have built and automated all sorts of cool stuff with the new Agentic model. Now, if you want to know exactly how to plug it in, need to use open router. You actually can't use Newportal despite the official updates. You can see here the Technium, the co-founder of Hermes, actually said the new Deep Seek V4 Flash is now available in Hermes agent through news portal and open router. and then he replied like 6 hours later saying never [0:30] mind news this portal doesn't allow models that will train on…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Julian Goldie SEO
Agent OS Q&A: Setup, Loops + New AI Models
A Q&A session covering the Agent OS—a modular AI automation system where users can plug in various AI models and tools. The speaker demonstrates setup processes, discusses token optimization techniques, explains looping features in Claude Code, and addresses community questions about integrating multiple AI agents without overwhelming complexity.
Claude AI SEO Gets Me 2,680 Clicks a Day
The speaker demonstrates how to generate 2,680 clicks per day using an AI-powered SEO system that leverages social media content ranking on Google. The strategy involves using Google Search Console data to identify target keywords, then creating automated video content across multiple platforms to rank multiple times for the same keyword.
Antigravity 2.0 + Gemini 3.6 Flash is INSANE!
Google has upgraded its Antigravity AI agent platform with Gemini 3.6 Flash, a faster and more token-efficient model that enables multi-agent workflows. The system allows a single prompt to be split across multiple specialized agents (planner, writer, reviewer, critic, auditor) that work in sync through Agent OS, a shared memory layer that maintains consistent brand voice and context across all agents.
Impeccable Makes Claude, Codex and Kimi 10X Better
Impeccable is a free, open-source design tool that improves AI-generated web designs by eliminating generic templates and providing design direction through product/design files and 23 specific commands. The tool works across Claude, Codex, and other models, detecting and fixing common AI design flaws like repetitive gradients, poor spacing, and accessibility issues before deployment.
NEW DeepSeek V4 Flash Update!
DeepSeek V4 Flash has been officially released with significant performance improvements over its preview version, designed specifically for AI agents. The speaker demonstrates practical applications by building 50+ projects and integrating it into their agent operating system, while emphasizing it's not frontier-level but offers excellent speed and cost efficiency.