TechnicalInsightful

Luo Fuli: OpenClaw, Agent Frameworks — The AI Paradigm Has Already Changed Dramatically!

Zhang Xiaojun Podcast

Luo Fuli, head of Xiaomi's large model division, describes how her firsthand experience with OpenClaw over Spring Festival 2026 fundamentally changed her understanding of AI agent frameworks as a paradigm shift — not just a product. She explains how OpenClaw's open-source, sophisticated context orchestration enabled her team to dramatically accelerate research and model training, and outlines how this Agent era demands a new approach to model architecture, post-training, and organizational design.

Summary

Luo Fuli, responsible for Xiaomi's large language model team, gives a detailed account of how the Agent paradigm has fundamentally shifted since her intensive use of OpenClaw during the 2026 Spring Festival period. She initially dismissed OpenClaw as a glorified UI wrapper on top of Claude Code, but three consecutive nights of deep engagement — first discovering its emotionally intelligent product design, then using it for team management strategy, and finally applying it to research tasks like building User Agents for post-training — transformed her view entirely. She now considers OpenClaw a 'epoch-defining agent framework' rather than just a product.

Luo explains that OpenClaw's key differentiator lies in its meticulous context orchestration: layered persistent memory, multi-model dispatch (automatically routing to better models for specific weaknesses like video understanding), and a fully open-source architecture that allows users to rewrite memory systems, multi-agent logic, and workflow designs. She contrasts this with Claude Code, which is optimized for software engineering and is a black box. She argues that a well-designed Agent framework can compensate for significant model capability gaps — even enabling a small 3B model to perform tasks she considered impossible for it.

She describes how she forced her entire team to use OpenClaw after the holiday, with group chats generating collective intelligence that rapidly improved both the framework and the team's imagination of what the technology could accomplish. Within three to four weeks, they accomplished research milestones she estimates would have previously taken thirty to forty weeks. The key insight was that Agent frameworks enable parallel research workflows — ten ideas can run simultaneously across sub-agents rather than sequentially.

On the technical side, Luo goes deep on the architectural decisions behind MiMo V2 Flash and Pro. Both models use a Hybrid Attention structure designed primarily around long-context efficiency, incorporating Sliding Window Attention at a 7:1 ratio with Full Attention in Pro (up from 5:1 in Flash), plus Multi-Token Prediction (MTP) at inference time. She argues MLA (used in DeepSeek, GLM, Kimi) is poorly suited for the Agent era because it leaves no computational headroom for MTP acceleration and was designed under assumptions — short post-training cycles and fixed inference hardware — that no longer hold. She claims this Hybrid architecture achieves 80–150 tokens per second at competitive cost, making it naturally suited for long-context Agent workloads.

Luo discusses the V2 model family: Pro handles complex reasoning and Agent orchestration, Omni addresses multi-modal perception (including joint audio-video understanding), and TTS uses a novel discrete tokenization approach inspired by NLP to achieve strong generalization from limited style training data. She reveals the team is pursuing a unified LLM-style architecture for audio and experimenting with the same for images, motivated by architectural elegance and infrastructure unification, though she acknowledges this is a difficult research bet that Agent-assisted coding has somewhat reduced the urgency of.

On AGI timelines, Luo states she believed AGI was at least two years away just two months prior, but now estimates it within two years, putting current progress at approximately 20% and expecting to reach 60–70% by year's end. She identifies AI training AI — the model reaching the intelligence level of the top researchers who train it and then iterating on itself — as the pivotal milestone, which she believes is likely within one to two years.

She characterizes the competitive landscape as: pre-training gaps between top Chinese and US labs are essentially closed; the 1T+ parameter base model is the 'entry ticket' to compete at Claude Opus 4.6 levels; and the real race is now in Agent post-training RL scaling, which very few teams have actually executed at pre-training compute scales. She believes Chinese labs have structural architecture advantages but need to rapidly build out Agent RL infrastructure and post-training pipelines.

Organizationally, Luo runs a team of roughly 100 people (including interns) with no formal group structure, no hierarchy, and no fixed deadlines — operated as an internal startup within Xiaomi. She believes flat structures, cross-pollination between pre-training and post-training roles, and passion-driven management produce more creative and adaptive research than traditional team segmentation. She is increasingly recruiting sophomore and junior undergraduates for their cognitive flexibility and openness to new paradigms.

Key Insights

  • Luo Fuli argues that OpenClaw's true breakthrough is not its UI design but its meticulous context orchestration — including layered persistent memory, automatic multi-model dispatch to compensate for individual model weaknesses, and full open-source modifiability — which together allow it to compensate for significant model capability gaps. She reports that even a 3B model performed tasks she considered impossible for it when embedded in OpenClaw's framework.
  • Luo argues that MLA (Multi-head Latent Attention, used by DeepSeek, GLM, and Kimi) is poorly suited for the Agent era because it was designed under now-obsolete assumptions: short post-training cycles and fixed inference hardware. MLA leaves no computational headroom for MTP-based inference acceleration, making models slower and more expensive for long-context Agent workloads compared to Hybrid Attention architectures like MiMo V2.
  • Luo claims that in the Agent paradigm, post-training compute should equal pre-training compute — a ratio of approximately 1:1 — and that research compute should exceed both at a ratio of roughly 3:1:1 (research : pre-train : post-train). She states this is a dramatic shift from the Chat era ratio she describes as roughly 1:5:1 in favor of pre-training.
  • Luo describes how she forced her team to use OpenClaw after Spring Festival by declaring anyone with fewer than 100 conversation turns the next day could quit — while privately having no intention of actually enforcing this. The real goal was forcing experiential exposure, because she believes experiencing a technology firsthand is the most effective way to ignite passion and imagination in a team, and that collective group experimentation multiplies individual imagination.
  • Luo argues that previous Agent frameworks and benchmarks (SWE-bench, BrowseComp, TAO-Bench) were fundamentally too simple and task-specific to constitute real Agent capability — they were essentially Chat with a slightly more complex system prompt and minimal environmental feedback. She states her team abandoned all such benchmarks when training MiMo V2, relying instead on body-sense evaluation within complex real Agent frameworks like Claude Code and OpenClaw as the true test of industrial-grade usability.

Topics

OpenClaw as an epoch-defining Agent frameworkMiMo V2 model architecture: Hybrid Attention, MTP, long-context efficiencyAgent post-training RL scaling as the new competitive frontierMulti-modal model development (Omni, TTS, discrete audio tokenization)Organizational design for AI research: flat structure, passion-driven management, no deadlinesAGI timeline and self-improving AI systemsChina vs. US AI competitive dynamicsOpenClaw vs. Claude Code framework comparisonCollective intelligence and open-source Agent framework development

Transcript

[0:01] Hello 大家好 我是小珺 今天我们的嘉宾是罗福莉 媒体叫她AI天才少女 但她不喜欢这个称呼 她目前是小米大模型的负责人 访谈是在OpenClaw发布之后 也是在2026年小米MiMo V2的系列模型发布之后 我们更深入地聊了聊 2026年由OpenClaw引发的新一轮的技术范式的 变迁以及未来技术演进的前沿话题 接下来就是我对罗福莉的访谈 [0:31] 期待2026年 我们和AI共同进步 这些能力都是可以被 我觉得最多一两个月 慢的话 三四个月 确实都可以被快速习得 所以环境反而比经验更重要 你刚才也提到1T的模型可能 是未来竞争的一个入场券 是这样吗 是Agent你要做到接近Claude Opus 4.6的 水平的这样一个入场券 那我如果说我们这样子来说 [1:02] 就是for研究 for Pre Train和for Post Train 对 我自己觉得一个非常合理的卡的 一个比例是可能3:1:1 对 就Pre Train和Post Train应该比例是 投入的算力是相当的 然后研究的比例应该至少是你 正式起训练的卡总量的还要多一点 就你要额外留更多的卡来去做研究 你过年的时候有跟我说 [1:33] 你觉得技术这几个月其实已经变天了 能不能阐述一下你觉得 过去两个月的这个技术的突变 我觉得 一个非常大的一个分界点 在于使用OpenClaw的前后 嗯 我自己其实是会把OpenClaw把它当做 一个划时代的Agent框架去这么去定义 嗯 嗯 我知道很多人在 尤其是用Claude Code做严肃编码的人 就会觉得OK OpenClaw是Claude Code加一个IM(即时通信)的 [2:03] 这样的一个更有利于交互的一个 UI的一个设计 其实在我1月份的时候 我第一次看到这个东西的时候 我自己大概也是这样认知 所以 我很排斥去用它 然后我觉得 而再加上创始人 我觉得非常适合贴近Agent去做 一些非常玄幻的一些运营的动作 所以就包括那个Skillhub啊这些的 让你更去排斥去用 一个你觉得非常的 [2:36] 偏运营导向的一个产品的东西 更感觉它是一个产品 一个交互范式 是对一个交互的创新 以及它所谓的本地化 所谓的24小时 在我来看 其实都是 一些产品的定义而已 但真正发生一个转变 是我去用它的那一刻 我觉得就恰好在 春节的时候 有那么一段空闲的时间 你想去搞明白这玩意 为什么它们那么火 对然后我就在有一天 深夜的时候去尝试去装了它 然后两个小时装上了 [3:07] 春节是吧 对当时已经凌晨2点了 然后我第一次跟它对话的时候 从凌晨2点持续到了6点天亮 对就我那一晚上我 觉得我脑内的那个 不知道是多巴胺还是内啡肽 就持续在分泌 就是让我就兴奋到完全睡不着觉 就你可能第一个感受是OK 它非常有自主性 然后它非常有灵魂 就比如说我跟它聊得很晚 它会老提醒我OK 你现在已经很晚了…

Full transcript available for MurmurCast members

Sign Up to Access

More from Zhang Xiaojun Podcast

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.