Qwen 3.8 Max, DeepSeek e Kimi: cosa sta succedendo davvero nell'AI
Chinese AI laboratories are releasing frontier open-source models weekly at 1/3 to 1/4 the cost of American models, forcing dramatic token price depreciation and shifting the AI industry's competitive focus from model intelligence (now a commodity) to hardware infrastructure, chips, and software optimization. This geopolitical competition is reshaping business models across the industry, with major companies pivoting toward open-source releases and custom hardware development.
Summary
The transcript presents a comprehensive analysis of the current state of AI competition, particularly between Chinese and American companies. Chinese laboratories including DeepSeek, Alibaba (Qwen), and others are releasing frontier-level open-source models approximately weekly, achieving performance parity with American models like GPT-4o at significantly lower costs—DeepSeek V4 Flash costs 8 cents per million output tokens versus GPT-4o Luna's $1.20. This price competition has forced OpenAI to depreciate token costs by 80% and is driving a fundamental restructuring of the AI industry's business model. The speaker argues that Large Language Models have become commodities, meaning performance alone no longer differentiates products. Consequently, the competitive battlefield has shifted from "who has the smartest model" to infrastructure, hardware efficiency, and software optimization (referred to as "harness"). Chinese companies maintain cost advantages because they develop their own chips, control manufacturing, and operate their own data centers without relying on U.S. semiconductor vendors. In response, American technology companies including Nvidia, Microsoft, Meta, and even startups like Mistral and Anthropic are increasingly releasing open-source models, recognizing that the real profit lies in hardware sales rather than token subscriptions. Nvidia's DJX Station and rumors of Apple's M7 Ultra with 1.5TB of unified memory represent attempts to capture the local inference market. The speaker presents financial data showing that major AI companies are currently unprofitable (losses ranging from billions to hundreds of billions), while hardware manufacturers like Micron and Nvidia are highly profitable, validating the thesis that hardware, not models, is where value accumulates. The analysis also addresses consumer hardware pricing increases, Apple's shift toward device leasing models, and speculation about future consolidation where companies will offer complete packages combining custom hardware, proprietary models, and software ecosystems. The speaker expresses concern about the "rental economy" approach to hardware ownership and discusses AI capability expansion, citing Alibaba's Qwen 3.8 Max achieving 16-day autonomous operation (versus previous single-day capabilities) and implications for prompt engineering transitioning to "loop engineering." The transcript concludes with predictions that model sizes will continue growing exponentially for data centers while simultaneously developing smaller, optimized models for edge devices and consumer hardware, fundamentally reshaping how AI infrastructure and business models evolve.
Key Insights
- Chinese laboratories release frontier-level open-source models approximately every week, achieving performance equivalent to American models at 1/3 to 1/4 the cost, forcing 80% depreciation in OpenAI's token pricing within weeks
- Large Language Models have become commodities where performance is no longer a differentiator—three competitors release equally intelligent models within the same week—shifting competition to hardware efficiency, custom scaffolding, and personnel quality
- Major AI companies including Amazon, Google, Meta, Microsoft, and OpenAI are currently unprofitable with combined losses in the hundreds of billions, while hardware manufacturers Nvidia and Micron are highly profitable, indicating the profit center has shifted from models to infrastructure
- Chinese companies maintain cost advantages by developing proprietary chips entirely made in China, controlling manufacturing facilities, and operating their own data centers without relying on U.S. semiconductors, enabling them to optimize architecture specifically for their models
- Alibaba's Qwen 3.8 Max can operate autonomously for 16 days continuously—compared to previous capabilities of hours or single days—fundamentally changing how AI work is orchestrated and requiring new engineering approaches like 'loop engineering' with precise termination conditions
Topics
Transcript
[0:00] We have Chinese laboratories releasing frontier open source models , practically one every week, which reach the performance of frontier American models and from the point of view of IP cost they cost 1/3 1/4, 1 deo compared to American models and they are trying in every way to destroy the business model of American companies which is based on selling you tokens to use the model. What did this entail? a [0:30] very significant depreciation in the cost of the token. If we look historically there is a significant depreciation. Lately Openi has depreciated by 80% the cost of PI on GPT 5.6 6 Luna and all this is due and caused by the Chinese who immediately after…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Simone Rizzo
gpt 5 6 sol, grock 4 5 e muse spark 1 1
A comprehensive review of recent AI model releases including Meta's Muse Spark 1.1, Elon Musk's Grock 4.5, and OpenAI's GPT 5.6 family (Sun, Earth, Moon variants). The speaker analyzes performance benchmarks, pricing, and capabilities, concluding that while new models are competitive, none have achieved a significant leap beyond current frontier models like Claude Fable 5.
Tutti parlano di Loop Engineering... ma nessuno te lo spiega così
The video traces the evolution of AI interaction paradigms from Prompt Engineering through Context Engineering, Harness Engineering, to the newest Loop Engineering approach. Loop Engineering involves wrapping autonomous AI workflows in iterative loops that self-improve toward defined goals without requiring manual intervention between steps.
DeepSeek ha appena reso TUTTI gli LLM più veloci
DeepSeek's new Spark technique uses semi-autoregressive speculative decoding to accelerate LLM inference by 51-400% without quality loss or model retraining. By combining a fast parallel draft model with an efficient verification process, Spark achieves higher token acceptance rates than competing methods like Eagle 3 and Flash, enabling faster inference on consumer hardware.
GLM 5.2 gira in locale quantizzandolo ad 1bit! #intelligenzaartificiale #aiagent
Researchers successfully ran the 744-billion parameter GLM 5.2 model locally on a Mac Studio M3 Ultra using dynamic quantization, compressing it from 810 GB to 223 GB. The 1-bit quantized version maintains 76.2% accuracy while being 86% smaller, and performs comparably to closed-source models like Claude Opus and GPT-5.5.