DeepSeek V4 Flash 0731 + Hermes Agent is INSANE!
DeepSeek V4 Flash, a newly released Chinese AI model optimized for agentic tasks, integrates with Hermes Agent to enable powerful automation workflows. The model offers significant performance improvements over its predecessor, cheaper pricing than frontier models, and a 1M token context window ideal for multi-step agent operations.
Summary
The video discusses DeepSeek V4 Flash (0731 build), a new model update that just became available through OpenRouter integration with Hermes Agent. The speaker clarifies that while initially announced for NewsPortal, the platform doesn't allow models trained on user data, so OpenRouter is the correct integration point. DeepSeek V4 Flash represents a major performance leap compared to the older Flash preview version, with superior agentic benchmarks. The model is cheaper than frontier alternatives like Claude Opus while maintaining comparable performance. The 0731 build is specifically a post-training upgrade focused on agent capability, featuring the same parameter size as before but with significantly better training for agent loops and tool calling. The speaker emphasizes that what previously limited agents wasn't Hermes itself, but the underlying model—factors like slow responses, poor tool-calling ability, memory loss over long jobs, and high frontier pricing all affected performance. DeepSeek V4 Flash addresses these through its 1M token context window, agentic-focused training, speed, and affordability. The speaker demonstrates three major use cases: creating AI avatar videos with full editing and voiceover, generating blog posts through multi-agent orchestration, and direct chat-based interactions. The content highlights the importance of building modular systems (Agent OS) rather than depending on specific models, as new models constantly emerge. The speaker promotes AI Profit Boardroom, a community offering the Agent OS, daily video tutorials, 24/7 support, and weekly calls for people building AI agent workflows.
Key Insights
- DeepSeek V4 Flash is specifically designed for agentic tasks with post-training focused on agent capability and tool-calling loops, making it fundamentally different from general-purpose frontier models despite being the same parameter size as the previous version
- The bottleneck for agent performance has historically been the underlying language model, not the Hermes agent framework itself—factors like model memory loss over 20+ steps, frontier pricing for each tool result, and lack of agentic training limited agent effectiveness
- DeepSeek V4 Flash's 1 million token context window is critical for multi-step agent jobs, allowing the model to remember everything throughout extended workflows without forgetting previous steps
- Building a modular Agent Operating System that can plug different models in and out is more valuable than optimizing for any single model, since new models constantly emerge and a flexible system ensures workflows never break
- DeepSeek V4 Flash response speed is significantly faster than frontier models like Claude through OpenRouter, delivering replies within seconds compared to 10-15 seconds, resulting in 2-3x more throughput in chat interactions
Topics
Transcript
[0:00] Deep Seek V4 Flash, the brand new update that just dropped yesterday, is now available on Hermes. I've tested I'm going to show you exactly how we have built and automated all sorts of cool stuff with the new Agentic model. Now, if you want to know exactly how to plug it in, need to use open router. You actually can't use Newportal despite the official updates. You can see here the Technium, the co-founder of Hermes, actually said the new Deep Seek V4 Flash is now available in Hermes agent through news portal and open router. and then he replied like 6 hours later saying never [0:30] mind news this portal doesn't allow models that will train on…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Julian Goldie SEO
How to Run DeepSeek V4 Flash for FREE!
A tutorial demonstrating how to use DeepSeek V4 Flash for free through Open Code and integration with agent operating systems like Hermes Agent. The speaker showcases building websites and apps using this free AI model and explains how it compares favorably to larger models despite being smaller.
Microsoft Fara1.5 27B NEW Browser Automation Model is WILD!
Microsoft released Phi-3.5, a family of three computer use models (4B, 9B, 27B) that automate browser tasks through vision-based clicking rather than HTML parsing. These open-weight models significantly outperform larger closed-source alternatives like OpenAI's Operator and Google's Gemini 2.0 on web automation benchmarks.
Claude Obsidian 2.0 is INSANE (FREE!)
Claude Obsidian 2.0 is presented as a free AI memory upgrade that allows users to upload files into a folder for permanent retention and linking. The system creates a knowledge graph that learns from business documents, provides sourced answers, and can be shared across teams.
Claude Agent OS is INSANE! 🤯
Julian presents a comprehensive Claude-based agent operating system that integrates multiple AI models, automated workflows, and a persistent memory system to automate daily tasks. The system runs 24/7 and uses free or existing subscriptions, combining tools like voice agents, content creation, competitor monitoring, and real-time news analysis into a single unified dashboard.
NEW ChatGPT Update is INSANE!
OpenAI released a major ChatGPT update featuring a Chrome extension called Side Chat and an improved desktop app that work together to streamline SEO research and content creation. The update allows users to analyze multiple browser tabs simultaneously, highlight text for quick answers, and convert research into finished work without constant tab switching.