Usare Jev per risparmiare tantissimi token nel tuo agente AI mantenendo alte le prestazioni.
Tamara discovered a method to use JEV, a fast decision-making model, to replace the slow slash compact function in AI agents. By having JEV decide whether to keep or discard function outputs, the system clears context window space much faster while maintaining performance. The solution, called Fast Jeev Comp, has gained significant popularity and can also be run locally using open-source models.
Summary
The transcript discusses a novel approach to optimizing token usage in AI agents. JEV is introduced as a specialized model with a particular architecture designed for speed, with the specific capability of making decisions and returning Boolean values, scores, or choices. A common problem in agent systems is that as conversations progress, the context window fills up, requiring a cleanup function called slash compact. However, slash compact is slow because it creates summaries of previous context. Tamara's innovation addresses this inefficiency by leveraging JEV's speed to make simple binary decisions about whether function outputs should be retained or discarded. This approach frees up space much more quickly than traditional summarization methods. The solution has been implemented as a Cloud Code Plugin called Fast Jeev Comp, which has achieved significant adoption with over 6,800 GitHub stars. The transcript also notes that the same approach can be replicated locally using open-source models with similar functionality to JEV, such as Laya or Rizzo Flow models, making the solution more accessible and flexible for different deployment scenarios.
Key Insights
- JEV is specifically designed with a particular architecture that prioritizes speed and has the single function of making decisions by responding with Boolean values, scores, or choices
- As AI agent conversations progress, the context window fills up and requires cleanup using slash compact, which is slow because it creates summaries of previous context
- Tamara's approach uses JEV to make fast binary decisions about whether to keep or discard function outputs, rather than creating summaries, which frees up space much more quickly
- Fast Jeev Comp, the implementation of this solution, has become very popular as a Cloud Code Plugin with over 6,800 GitHub stars
- The same optimization pattern can be implemented locally using open-source models like Laya or Rizzo Flow models instead of relying on proprietary JEV
Topics
Transcript
[0:00] This is Tamara and she found a way to use Jeev to save you tons of tokens. For those who don't know, JEV is this new model that has this particular architecture that is super fast, but it only does one thing, which is it makes decisions and responds with a Bulean value or a score or choices. When using any agent you will find that as the conversation progresses, the context window starts to fill up. Once you get to a certain point there is a function called slash compact. This is what [0:31] makes you try to clean up all the previous context by freeing up space. Now Slash Compact works but it is very slow, it…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Simone Rizzo
OpenAI dots compete con grokbot il nuovo gpt 6.1 sol compete con opus 5.5 tutto questo è successo
OpenAI released multiple new frontier models including GPT 6.1 Sol at significantly lower token costs, along with cloud-based agents called Dots that compete with Grockbot. The company introduced new subscription tiers, collaborative spaces, a Decisions API for real-time decision-making, and cloud-based coding tools, while Elon Musk strategically acquired dots.com to redirect traffic to Grockbot.
OpenAI copia Grok con i dots? GPT-6.1 sfida Opus 5.5
A comprehensive analysis of OpenAI's DevDay announcements, including new AI agents called 'Dots,' the GPT-6.1 model, and pricing changes, alongside broader trends in AI development from competitors like Anthropic and open-source projects. The speaker critiques the announcements as largely incremental improvements rather than groundbreaking innovations.
Ti svelo il miglior modello AI locale da far girare sul tuo pc con 4Miliardi di parametri.
Spark X 2.5 is highlighted as the best 4-billion parameter local AI model offering an optimal balance between speed and performance. The model outperforms competitors in its size class, matches the performance of 9B models, and can handle up to 1 million input tokens, making it suitable for various local AI applications.
Jev è a pagamento? L'ho rifatto open source e te lo regalo
The creator explains how Jamba (an AI decision-making model) works and presents Rizzo Flow, an open-source alternative built in just two days. The video demonstrates that recreating Jamba is straightforward once the architecture is understood, using either encoders or decoders to make fast, typed decisions without generating tokens.
Qwen 3.8 Max, DeepSeek e Kimi: cosa sta succedendo davvero nell'AI
Chinese AI laboratories are releasing frontier open-source models weekly at 1/3 to 1/4 the cost of American models, forcing dramatic token price depreciation and shifting the AI industry's competitive focus from model intelligence (now a commodity) to hardware infrastructure, chips, and software optimization. This geopolitical competition is reshaping business models across the industry, with major companies pivoting toward open-source releases and custom hardware development.