TechnicalInsightful

Jev è a pagamento? L'ho rifatto open source e te lo regalo

Simone Rizzo

The creator explains how Jamba (an AI decision-making model) works and presents Rizzo Flow, an open-source alternative built in just two days. The video demonstrates that recreating Jamba is straightforward once the architecture is understood, using either encoders or decoders to make fast, typed decisions without generating tokens.

Summary

The video begins by explaining that Jamba's success comes from being a 'System One' model—a fast decision-making system rather than a traditional language model. It makes parallel typed responses (boolean, class, or score) in milliseconds without generating tokens. After Jamba's release, the open-source community quickly created multiple alternatives.

The creator discusses three architectural approaches to recreating Jamba: encoders, decoders, and diffusion models. Encoder-based models like Laya use bidirectional attention for classification tasks but require custom training and output layers. Decoder-based models (the majority of implementations) are easier to adapt because they already generate predictions—you simply remove autoregressive token generation and extract logits from the final layer. By examining probabilities of specific tokens representing classes or numbers, the model can instantly provide decisions without sequential token output.

The creator introduces Rizzo Flow, their open-source implementation built on Spark 2.5 decoder with 4 billion parameters. Key advantages include support for 1 million input tokens (versus Jamba's 32,000 and Semif's 262), four output primitives instead of two, and up to 26 multiple-choice options per question. The model was fine-tuned using LoRA on distilled Jamba data to calibrate probability distributions, improving performance metrics like KL divergence and Brier score.

Rizzo Flow uses KV-cache optimization to process multiple questions efficiently by reusing cached state representations and only recalculating attention for new questions. The implementation maintains the same API as Jamba, runs locally on any hardware (CPU, GPU, Mac, Linux), and can be deployed via Docker. The creator addresses misconceptions: distilling data from Jamba isn't illegal (it's a terms-of-service issue, not legal); the two-day turnaround is possible because the architectural 'recipe' is now public; and the model's value lies in solving practical business problems rather than architectural novelty.

Key Insights

  • Decoders can function as System One models by removing autoregressive generation and examining only the logits of relevant output tokens in the final layer, enabling instant typed responses instead of sequential token generation
  • KV-cache optimization allows multiple questions to be answered in parallel by pre-computing and reusing the key-value matrices for the fixed state input, then only recalculating attention for each new question
  • Rizzo Flow supports 1 million input tokens natively compared to Jamba's 32,000 and Semif's 262, enabling longer contextual reasoning and more complex automation horizons
  • Distilling data from Jamba by saving user queries and responses is not illegal but violates the company's terms of service—the worst consequence is profile banning, not legal action
  • Once an architectural recipe is publicly demonstrated, rebuilding models like Jamba becomes straightforward because the hard problem (discovering the approach) is already solved, similar to how Transformers were replicated quickly after their paper release

Topics

System One models vs large language modelsEncoder vs decoder architectures for decision-makingLogits extraction and probability calibration techniquesKV-cache optimization for multi-question inferenceOpen-source model distillation and fine-tuningRizzo Flow implementation and featuresReal-world use cases for fast classifiers

Transcript

[0:00] Jeev is the model of the moment, but it is closed and paid. So I decided to take a stab at it and create an open source version of Jeev that I called Rizzo Flow. But let's proceed in order. Jeev has been so successful because it is not a large language model, it is not a chatbot, it does not produce tokens, but it is the first system-one model, that is, a model that makes decisions at the speed of light. Specifically, we can see it as an [0:30] intelligent function call that receives more state than input questions, answers those questions in parallel with typed answers, and then responds with a Bulean, yes, no, or a class…

Full transcript available for MurmurCast members

Sign Up to Access

More from Simone Rizzo

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.