Jev è a pagamento? L'ho rifatto open source e te lo regalo
The creator explains how Jamba (an AI decision-making model) works and presents Rizzo Flow, an open-source alternative built in just two days. The video demonstrates that recreating Jamba is straightforward once the architecture is understood, using either encoders or decoders to make fast, typed decisions without generating tokens.
Summary
The video begins by explaining that Jamba's success comes from being a 'System One' model—a fast decision-making system rather than a traditional language model. It makes parallel typed responses (boolean, class, or score) in milliseconds without generating tokens. After Jamba's release, the open-source community quickly created multiple alternatives.
The creator discusses three architectural approaches to recreating Jamba: encoders, decoders, and diffusion models. Encoder-based models like Laya use bidirectional attention for classification tasks but require custom training and output layers. Decoder-based models (the majority of implementations) are easier to adapt because they already generate predictions—you simply remove autoregressive token generation and extract logits from the final layer. By examining probabilities of specific tokens representing classes or numbers, the model can instantly provide decisions without sequential token output.
The creator introduces Rizzo Flow, their open-source implementation built on Spark 2.5 decoder with 4 billion parameters. Key advantages include support for 1 million input tokens (versus Jamba's 32,000 and Semif's 262), four output primitives instead of two, and up to 26 multiple-choice options per question. The model was fine-tuned using LoRA on distilled Jamba data to calibrate probability distributions, improving performance metrics like KL divergence and Brier score.
Rizzo Flow uses KV-cache optimization to process multiple questions efficiently by reusing cached state representations and only recalculating attention for new questions. The implementation maintains the same API as Jamba, runs locally on any hardware (CPU, GPU, Mac, Linux), and can be deployed via Docker. The creator addresses misconceptions: distilling data from Jamba isn't illegal (it's a terms-of-service issue, not legal); the two-day turnaround is possible because the architectural 'recipe' is now public; and the model's value lies in solving practical business problems rather than architectural novelty.
Key Insights
- Decoders can function as System One models by removing autoregressive generation and examining only the logits of relevant output tokens in the final layer, enabling instant typed responses instead of sequential token generation
- KV-cache optimization allows multiple questions to be answered in parallel by pre-computing and reusing the key-value matrices for the fixed state input, then only recalculating attention for each new question
- Rizzo Flow supports 1 million input tokens natively compared to Jamba's 32,000 and Semif's 262, enabling longer contextual reasoning and more complex automation horizons
- Distilling data from Jamba by saving user queries and responses is not illegal but violates the company's terms of service—the worst consequence is profile banning, not legal action
- Once an architectural recipe is publicly demonstrated, rebuilding models like Jamba becomes straightforward because the hard problem (discovering the approach) is already solved, similar to how Transformers were replicated quickly after their paper release
Topics
Transcript
[0:00] Jeev is the model of the moment, but it is closed and paid. So I decided to take a stab at it and create an open source version of Jeev that I called Rizzo Flow. But let's proceed in order. Jeev has been so successful because it is not a large language model, it is not a chatbot, it does not produce tokens, but it is the first system-one model, that is, a model that makes decisions at the speed of light. Specifically, we can see it as an [0:30] intelligent function call that receives more state than input questions, answers those questions in parallel with typed answers, and then responds with a Bulean, yes, no, or a class…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Simone Rizzo
OpenAI dots compete con grokbot il nuovo gpt 6.1 sol compete con opus 5.5 tutto questo è successo
OpenAI released multiple new frontier models including GPT 6.1 Sol at significantly lower token costs, along with cloud-based agents called Dots that compete with Grockbot. The company introduced new subscription tiers, collaborative spaces, a Decisions API for real-time decision-making, and cloud-based coding tools, while Elon Musk strategically acquired dots.com to redirect traffic to Grockbot.
OpenAI copia Grok con i dots? GPT-6.1 sfida Opus 5.5
A comprehensive analysis of OpenAI's DevDay announcements, including new AI agents called 'Dots,' the GPT-6.1 model, and pricing changes, alongside broader trends in AI development from competitors like Anthropic and open-source projects. The speaker critiques the announcements as largely incremental improvements rather than groundbreaking innovations.
Usare Jev per risparmiare tantissimi token nel tuo agente AI mantenendo alte le prestazioni.
Tamara discovered a method to use JEV, a fast decision-making model, to replace the slow slash compact function in AI agents. By having JEV decide whether to keep or discard function outputs, the system clears context window space much faster while maintaining performance. The solution, called Fast Jeev Comp, has gained significant popularity and can also be run locally using open-source models.
Ti svelo il miglior modello AI locale da far girare sul tuo pc con 4Miliardi di parametri.
Spark X 2.5 is highlighted as the best 4-billion parameter local AI model offering an optimal balance between speed and performance. The model outperforms competitors in its size class, matches the performance of 9B models, and can handle up to 1 million input tokens, making it suitable for various local AI applications.
Qwen 3.8 Max, DeepSeek e Kimi: cosa sta succedendo davvero nell'AI
Chinese AI laboratories are releasing frontier open-source models weekly at 1/3 to 1/4 the cost of American models, forcing dramatic token price depreciation and shifting the AI industry's competitive focus from model intelligence (now a commodity) to hardware infrastructure, chips, and software optimization. This geopolitical competition is reshaping business models across the industry, with major companies pivoting toward open-source releases and custom hardware development.