NewsTechnical

Apple rilascia un nuovo modello AI si chiama LensVLM e cerca di risolvere il problema dei token in

Simone Rizzo

Apple releases LensVLM, a fine-tuned 9 billion parameter model based on Qwen 3.5, implementing a novel technique called Selective Context Expansion that converts text into images to achieve 15x token compression. This approach allows visual language models to handle large text inputs more efficiently while maintaining reasoning accuracy.

Summary

Apple, traditionally focused on hardware rather than AI model development, has released a new visual language model called LensVLM with 9 billion parameters. The model is built by fine-tuning Qwen 3.5 and is accompanied by a scientific paper titled 'Selective Context Expansion for Compressed Visual Representation of Text,' which introduces an innovative approach inspired by Deep Seek's OCR research. The key innovation addresses a fundamental challenge in large language models: handling extremely large text inputs that consume excessive tokens. Rather than processing text directly, the technique converts text into visual representations (images containing the text), which are then converted into vision tokens. This conversion achieves a 15x compression ratio because vision tokens consume significantly less space than text tokens. Once compressed and input into the model, the system can selectively expand and reason about specific parts of the image to accurately understand concepts and generate appropriate responses. The speaker expresses surprise at Apple's entry into model research, noting that the company has intelligently positioned itself primarily in the hardware space, where it manufactures powerful devices like Mac Studio with M5 and M6 chips capable of running cutting-edge models like Deep Seek. The speaker speculates that Apple may continue selective research and fine-tuning efforts, potentially releasing more papers and optimized models, though it remains uncertain whether Apple will develop a fully competitive open-source model to challenge Chinese and American competitors or maintain its focus on hardware.

Key Insights

  • Apple has intelligently avoided competing in the AI model development race, instead focusing on hardware manufacturing where the real profit margins exist, such as producing powerful Mac Studio devices with M5 and M6 chips
  • The Selective Context Expansion technique converts large text inputs into visual representations that render as images, which are then converted to vision tokens, achieving 15x token compression compared to text tokens
  • Deep Seek's OCR paper served as the primary inspiration for Apple's new compression technique, suggesting cross-pollination of ideas between different AI research groups
  • LensVLM maintains reasoning accuracy while maximally compressing the required tokens by selectively expanding parts of compressed visual representations to process specific concepts
  • Apple's strategy appears to involve selective research efforts, publishing scientific papers and releasing fine-tuned models, rather than committing to developing a fully independent large-scale open-source model

Topics

Apple LensVLM model releaseSelective Context Expansion techniqueToken compression for visual language modelsText-to-image conversion for efficiencyApple's AI strategy and hardware focus

Transcript

[0:00] And this one, I really missed this one. Apple releasing a new model? Apple, which intelligently, in my opinion , did not participate in the race to develop AI models, but focused on hardware, where the real money is. And in fact it releases these super-powerful Mac Studio M5, M6 that let you run models like Deep Seek, which are cutting-edge models for everyone. And so I was surprised by this new news about this model called Lens VLM 9 [0:30] billion, which is nothing more than taking the 9 billion Qwen 3.5 and doing some fine tuning on it. So I got curious and said, “But why are they doing this?” And then I saw that associated with…

Full transcript available for MurmurCast members

Sign Up to Access

More from Simone Rizzo

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.