Fei Fei Li: The Race to Build World Models For AI
World Labs launched Atlas, a new world model built on next-view prediction that unifies 3D reconstruction and generation. The model can create spatially-grounded video frames from sparse input images (as few as three), reducing the data requirements for 3D scene capture by 50-100x, with applications ranging from creative content to robotics simulation.
Summary
Fei-Fei Li, Justin Johnson, and Ben Mildenhall from World Labs discussed their new frontier model Atlas, which represents a fundamental shift in how AI approaches spatial intelligence. Unlike previous models built on next-token prediction (LLMs) or next-frame prediction (video models), Atlas is built on next-view prediction—given some views of a scene, it predicts what the scene looks like from any arbitrary point in space and time.
Atlas combines two traditionally separate computer vision tasks: dense 3D reconstruction and generative pixel synthesis. Previously, reconstruction required exhaustive capture of a scene with hundreds of images to ensure every angle was covered, making it impractical for most applications. Atlas reduces this requirement dramatically—it can reconstruct complex scenes from as few as three images, representing a 50-100x reduction in capture requirements. This is achieved by combining triangulation from multiple views (reconstruction) with generative filling of occluded regions (generation).
The model is multimodal by design, natively working with text, images, videos, camera poses, and depth maps as inputs. Camera poses are a critical native input—every frame has an associated 3D camera position and parameters, allowing spatial grounding that distinguishes Atlas from other video generation models that lack this geometric precision.
Key capabilities include camera-conditioned generation (steering viewpoints through space), sparse 3D reconstruction (turning multi-view captures into navigable 3D scenes), and simulation including the famous bullet-time effects from The Matrix. Traditionally requiring hundreds of cameras and green screen studios, these effects now work with just three iPhones on tripods.
The team explained that achieving this required moving beyond their previous product Marble, which output Gaussian splats and couldn't gracefully scale to handle many input images or extensive context windows. Atlas bifurcates modalities earlier in the architecture and treats new-view prediction as the fundamental primitive, allowing it to output either RGB frames, 3D reconstructions, or Gaussian splats depending on the use case.
On scaling, the team reported being at the beginning of the scaling curve for this architectural approach, limited primarily by compute rather than data or architecture. Each increase in model size and training time showed significant improvements. Ben Mildenhall's conviction in scaling came from witnessing dense reconstruction improvements over three years prior to the company's founding.
For creative applications, Atlas extends workflows where artists compose multi-stage pipelines. Rather than requesting different viewpoints from multiple models and dealing with inconsistency, they can now ground generations in a persistent 3D world with precise viewpoint control. Use cases extend to architecture, design, and any field requiring virtual replication of physical spaces.
The robotics application is particularly significant. World Labs acquired Cynex to address the critical bottleneck in robotics: data collection for sim-to-real transfer. Dense reconstruction of robot environments is currently laborious and slow. Atlas dramatically accelerates the real-to-sim step, enabling faster training data generation. The vision extends beyond reconstruction to learned simulators that can predict how environments respond to actions, potentially enabling these neural simulators to become both prediction and planning engines.
The team acknowledged that dynamics (motion, change over time) is an important missing piece, though they noted Atlas already contains some dynamic capabilities from training data exposure. They emphasized that exposing models to dynamics during training actually helps static reconstruction by letting the model factor out temporal variation.
Finally, the team proposed that next-view prediction is 'AI complete'—analogous to next-token prediction for language models. Any intelligence task can be framed as predicting the next viewpoint, making this a fundamental primitive for spatial intelligence and potentially general intelligence itself.
About this episode
World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence. At the center of Atlas is what the team calls “new view prediction”: given images or views of a scene, the model predicts what that environment should look like from a different position in space and time. This brings generation and 3D reconstruction into the same model, and raises a broader question about whether predicting views could become a useful primitive for understanding the physical world. They discuss the technical bets behind the model, what it can and can’t yet capture, and the importance of dynamics, editability, and simulation as world models develop. The conversation also explores applications in creative work, architecture, and robotics, where Fei-Fei argues that one of today’s biggest constraints is access to real-world training data.
Key Insights
- Atlas pioneers next-view prediction as a foundational AI primitive, analogous to how next-token prediction enables language models, with the hypothesis that next-view prediction could be equally fundamental to spatial intelligence.
- The model unifies 50+ years of separate computer vision research by combining 3D reconstruction and generative pixel synthesis in a single architecture, enabling accurate 3D scene modeling from sparse inputs rather than requiring exhaustive multi-view capture.
- Camera pose as a native input modality—not just an output—is critical for spatial grounding and enables the model to maintain 3D consistency that other video generation models lack, distinguishing it from approaches that interpret images without geometric constraints.
- The 50-100x reduction in capture requirements transforms practical applications: bullet-time effects now require three iPhone cameras instead of hundreds of studio cameras and green screens, and robotics environments can be digitized from casual video rather than dense professional capture.
- The team demonstrated they are at the beginning of the scaling curve for this architecture, with compute being the primary bottleneck rather than data or architectural limitations, suggesting substantial performance improvements remain achievable with more training resources.
- Exposure to dynamics during training actually improves static reconstruction quality by allowing the model to learn to factor out temporal variation, suggesting that mixing dynamic and static data improves both capabilities despite apparent tensions.
- The acquisition of robotics company Cynex and integration with Atlas addresses a critical bottleneck in AI robotics: the extreme difficulty of collecting real-world training data for sim-to-real transfer, which previously required dense reconstruction workflows.
- The team conceptualizes next-view prediction as potentially 'AI complete'—any intelligence task can be framed as predicting the next viewpoint, suggesting this primitive may be as foundational to spatial and general intelligence as next-token prediction is to language.
Topics
Transcript
On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken. We know LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction. This is the real place where AI can actually unlock a ton of value for people and their process. We're saying like 50, 100x reduction. There was a famous shot in the first Matrix movie where Neo is like falling down. Exactly. They had hundreds of cameras viewing that angle on a green screen. On Atlas, we can do this with just three cameras. No studio capture, no green…
Full transcript available for MurmurCast members
Sign Up to AccessMore from The a16z Show
The $100B Niches Hiding Inside Payments
Max Levchin and Alex Rampell discuss 25+ years of payments innovation, tracing PayPal's origins through their journey building Affirm. They explore how the credit card remains the best payment interface ever created, why there are no niches smaller than $100B in payments, and how AI may finally be ready to reinvent the payment experience itself.
Inside Moderna’s Personalized Cancer Vaccine
Moderna and Merck announced positive Phase 3 results for an individualized mRNA cancer vaccine (Intesmiran) for melanoma, achieving 80% disease-free survival at five years compared to 60% with Keytruda alone. The breakthrough combines mRNA technology with personalization, sequencing each patient's tumor to identify mutations and create a patient-specific vaccine that teaches the immune system to recognize cancer cells.
Daniel Litt: The Mathematician's Guide to AI
Daniel Litt, a mathematician at University of Toronto, discusses AI's evolving capabilities in mathematics, distinguishing between what models can do well (applying known techniques, solving specific problems) versus where they fall short (developing intuition, building new theories, asking fundamental questions). He emphasizes that while AI results are impressive, the mathematics community must adapt its incentive structures to preserve human understanding and maintain cognitive diversity in mathematical research.
Gavin Baker: Why AI Demand Is Outrunning Compute Supply
Gavin Baker and David George discuss why AI demand is outrunning compute supply, examining the positive-sum nature of the AI market where frontier labs, open-source models, cloud providers, and chip companies can all win. They argue that despite concerns about bubbles, the economics of AI infrastructure show sub-one-year paybacks, massive supply constraints, and early adoption suggesting we're nowhere near peak demand.
Why a16z Launched the Machine Age Fund | Jen Kha
Andreessen Horowitz launched a $1.1 billion Machine Age Fund to invest in physical AI infrastructure—chips, networking, data centers, and robotics—addressing a massive supply-side bottleneck as AI demand accelerates globally. The fund represents venture capital's return to hardware after 30 years of focusing on software, with a16z seeing hardware pitches rise from near-zero to over 20% of all submissions.