Apple rilascia un nuovo modello AI si chiama LensVLM e cerca di risolvere il problema dei token in
Apple releases LensVLM, a fine-tuned 9 billion parameter model based on Qwen 3.5, implementing a novel technique called Selective Context Expansion that converts text into images to achieve 15x token compression. This approach allows visual language models to handle large text inputs more efficiently while maintaining reasoning accuracy.
Summary
Apple, traditionally focused on hardware rather than AI model development, has released a new visual language model called LensVLM with 9 billion parameters. The model is built by fine-tuning Qwen 3.5 and is accompanied by a scientific paper titled 'Selective Context Expansion for Compressed Visual Representation of Text,' which introduces an innovative approach inspired by Deep Seek's OCR research. The key innovation addresses a fundamental challenge in large language models: handling extremely large text inputs that consume excessive tokens. Rather than processing text directly, the technique converts text into visual representations (images containing the text), which are then converted into vision tokens. This conversion achieves a 15x compression ratio because vision tokens consume significantly less space than text tokens. Once compressed and input into the model, the system can selectively expand and reason about specific parts of the image to accurately understand concepts and generate appropriate responses. The speaker expresses surprise at Apple's entry into model research, noting that the company has intelligently positioned itself primarily in the hardware space, where it manufactures powerful devices like Mac Studio with M5 and M6 chips capable of running cutting-edge models like Deep Seek. The speaker speculates that Apple may continue selective research and fine-tuning efforts, potentially releasing more papers and optimized models, though it remains uncertain whether Apple will develop a fully competitive open-source model to challenge Chinese and American competitors or maintain its focus on hardware.
Key Insights
- Apple has intelligently avoided competing in the AI model development race, instead focusing on hardware manufacturing where the real profit margins exist, such as producing powerful Mac Studio devices with M5 and M6 chips
- The Selective Context Expansion technique converts large text inputs into visual representations that render as images, which are then converted to vision tokens, achieving 15x token compression compared to text tokens
- Deep Seek's OCR paper served as the primary inspiration for Apple's new compression technique, suggesting cross-pollination of ideas between different AI research groups
- LensVLM maintains reasoning accuracy while maximally compressing the required tokens by selectively expanding parts of compressed visual representations to process specific concepts
- Apple's strategy appears to involve selective research efforts, publishing scientific papers and releasing fine-tuned models, rather than committing to developing a fully independent large-scale open-source model
Topics
Transcript
[0:00] And this one, I really missed this one. Apple releasing a new model? Apple, which intelligently, in my opinion , did not participate in the race to develop AI models, but focused on hardware, where the real money is. And in fact it releases these super-powerful Mac Studio M5, M6 that let you run models like Deep Seek, which are cutting-edge models for everyone. And so I was surprised by this new news about this model called Lens VLM 9 [0:30] billion, which is nothing more than taking the 9 billion Qwen 3.5 and doing some fine tuning on it. So I got curious and said, “But why are they doing this?” And then I saw that associated with…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Simone Rizzo
Gemini 4 Argon é stato appena annunciato da Google e rappresenta a detta loro il nuovo modello AI
Google announces Gemini 4 Argon, a new AI model claimed to compete with OpenAI and Anthropic offerings, featuring strong performance across benchmarks and enhanced security against prompt injection attacks. The model will be released through a staged rollout starting with the Fairwind Program for selected companies, followed by API access for developers, and eventually public release.
OpenAI dots compete con grokbot il nuovo gpt 6.1 sol compete con opus 5.5 tutto questo è successo
OpenAI released multiple new frontier models including GPT 6.1 Sol at significantly lower token costs, along with cloud-based agents called Dots that compete with Grockbot. The company introduced new subscription tiers, collaborative spaces, a Decisions API for real-time decision-making, and cloud-based coding tools, while Elon Musk strategically acquired dots.com to redirect traffic to Grockbot.
OpenAI copia Grok con i dots? GPT-6.1 sfida Opus 5.5
A comprehensive analysis of OpenAI's DevDay announcements, including new AI agents called 'Dots,' the GPT-6.1 model, and pricing changes, alongside broader trends in AI development from competitors like Anthropic and open-source projects. The speaker critiques the announcements as largely incremental improvements rather than groundbreaking innovations.
Usare Jev per risparmiare tantissimi token nel tuo agente AI mantenendo alte le prestazioni.
Tamara discovered a method to use JEV, a fast decision-making model, to replace the slow slash compact function in AI agents. By having JEV decide whether to keep or discard function outputs, the system clears context window space much faster while maintaining performance. The solution, called Fast Jeev Comp, has gained significant popularity and can also be run locally using open-source models.
Ti svelo il miglior modello AI locale da far girare sul tuo pc con 4Miliardi di parametri.
Spark X 2.5 is highlighted as the best 4-billion parameter local AI model offering an optimal balance between speed and performance. The model outperforms competitors in its size class, matches the performance of 9B models, and can handle up to 1 million input tokens, making it suitable for various local AI applications.