The limits of AI scaling laws - NVIDIA CEO explains | Jensen Huang and Lex Fridman
Jensen Huang discusses NVIDIA's perspective on AI scaling laws across four dimensions—pre-training, post-training, test-time, and agentic scaling—while addressing historical blockers that were overcome and future challenges like power efficiency and supply chain complexity. He argues that intelligence scales primarily with compute, explains how agentic systems require tools and data access rather than replacing existing infrastructure, and emphasizes the importance of anticipating hardware needs years in advance through co-design.
Summary
Jensen Huang explains that NVIDIA continues to believe in scaling laws for AI, outlining four dimensions of scaling: pre-training, post-training, test-time (inference), and agentic scaling. He reflects on historical blockers that industry experts predicted would limit AI progress. The first major blocker was data scarcity for pre-training, which some predicted would end AI scaling. However, Huang argues this concern was overcome through synthetic data generation, noting that most human knowledge is already 'synthetic' (created, modified, and regenerated by humans). He explains that training is now limited by compute rather than data, since most training data is synthetically generated.
Regarding test-time scaling and inference, Huang refutes the assumption that inference would be computationally simple compared to pre-training. He argues that inference involves thinking, reasoning, planning, and search—processes far more computationally intensive than the pattern-matching of pre-training. This scaling law, he contends, has proven correct: inference requires massive compute.
The fourth scaling law is agentic scaling, which Huang describes as 'multiplying AI.' Agents spawn sub-agents to accomplish tasks, creating teams of AI systems. He explains that agentic systems will access ground truth (file systems), conduct research, use tools, and communicate externally. These systems create enormous amounts of data and experiences that feed back into pre-training, creating a continuous improvement cycle.
Huang emphasizes the importance of anticipating hardware and system architecture needs 2-3 years in advance, since AI model architectures change every 6 months but hardware architectures take 3 years to develop. NVIDIA's strategy involves internal research, continuous dialogue with industry partners, and maintaining CUDA's balance between specialization and flexibility. He provides concrete examples: when mixture-of-experts emerged, NVIDIA anticipated this and developed NVLink-72 to handle massive models. Similarly, the Vera Rubin rack was designed to support agentic systems with tools, storage, and processing capabilities—a shift from the previous Grace-Blackwell architecture focused on inference.
Huang explains that he reasoned about agentic system requirements first-principles, drawing parallels to how a humanoid robot would use existing tools rather than transforming its hand into different tools. This reasoning led NVIDIA to design systems that support tool access, file systems, research capabilities, and external communication—properties that emerged in Claude's OpenClaw.
On supply chain challenges, Huang emphasizes that he actively works with upstream and downstream partners to shape the future. He describes convincing DRAM manufacturers to invest in HBM memory when its use was marginal, explaining why it would become mainstream. He also worked with manufacturers to adopt LPDDR5 memory (originally for cell phones) in supercomputers. The shift to building supercomputers in the supply chain rather than assembling them in data centers required partners to invest billions in manufacturing capability.
Regarding power consumption, Huang identifies energy efficiency (tokens per second per watt) as critical. NVIDIA has achieved a million-fold improvement in computing over 10 years compared to Moore's Law's 100-fold improvement. However, he emphasizes that absolute power availability is also necessary, proposing solutions involving grid dynamics and flexible power contracts.
Huang argues that power grids are designed for worst-case conditions but operate at 60% capacity 99% of the time. He proposes that data centers could gracefully degrade performance during peak grid demand, shifting workloads or accepting slightly longer latencies. He identifies three barriers: customer contracts demanding 'six nines' uptime, data centers not designed to gracefully degrade, and utilities not offering tiered power delivery guarantees. He suggests that if these three components align, data centers could use excess grid capacity efficiently without requiring massive new grid infrastructure.
Key Insights
- Jensen Huang argues that synthetic data will overcome the apparent blocker of limited high-quality training data, because most human knowledge is already 'synthetic' (created, modified, and shared by people), enabling continuous scaling of training data as long as compute is available.
- Huang contends that inference is fundamentally about thinking and reasoning rather than memorization, making it far more computationally demanding than pre-training, contrary to industry assumptions that inference would be 'easy' and 'compute light.'
- NVIDIA anticipates hardware needs 2-3 years in advance through internal research, industry listening, and maintaining CUDA's balance between specialization and generalization, enabling the company to stay ahead of rapidly evolving AI architectures.
- Huang reasoned first-principles about agentic systems two years before OpenClaw's public release, concluding they would need tool access, file systems, research capabilities, and communication infrastructure—similar to how a humanoid robot would use existing tools rather than replace them.
- Huang proposes that power grid inefficiency (99% of capacity sitting idle outside peak demand) can be addressed by redesigning data centers to gracefully degrade performance during peak grid demand, requiring changes to customer contracts, data center architecture, and utility pricing models.
Topics
Transcript
[0:02] - Yeah. So one of the things you've been a believer for a long time is scaling laws, broadly defined. So are you still a believer in the scaling laws? - Yeah, yeah. Yeah, we have more scaling laws now. - So I think you've outlined four of them with pre-training, post-training, test time, and agentic scaling. What do you think, when you think about the future, deep future and the near-term future, what are the blockers that you're most concerned about that keep you up at night that you have to overcome [0:33] in order to keep scaling? - Well, we can go back and reflect on what people thought were blockers. So in the beginning, we were…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Lex Clips
How Sigmund Freud revolutionized psychiatry | Andrew Scull and Lex Fridman
Andrew Scull discusses how Sigmund Freud and psychoanalysis emerged in late 19th-century America alongside religious healing movements, eventually gaining popularity among intellectuals and artists in the 1920s, despite mainstream psychiatry's resistance to talk therapy as treatment for mental illness.
Do antipsychotic drugs work? | Andrew Scull and Lex Fridman
Andrew Scull discusses the efficacy and serious side effects of antipsychotic drugs, explaining that while they help some patients, many are non-responders, and the medications carry substantial risks including tardive dyskinesia, weight gain, metabolic syndrome, and cognitive dulling. The landmark CATIE study revealed that newer, expensive second-generation antipsychotics are no more effective than older first-generation drugs, with 67-82% of patients dropping out due to inefficacy or intolerable side effects.
Surprising origin of antipsychotic drugs | Andrew Scull and Lex Fridman
Andrew Scull explains how chlorpromazine, an antihistamine synthesized in the 1880s, was accidentally discovered to have psychiatric applications in the 1950s through French Navy lieutenant Alain Laborit's experiments. The drug's rapid adoption was driven by pharmaceutical companies' aggressive marketing to politicians and hospital administrators rather than by psychiatrists' initial enthusiasm, transforming it from a 'major tranquilizer' into the first 'antipsychotic' drug.
Cognitive Behavioral Therapy (CBT) vs Psychoanalysis | Andrew Scull and Lex Fridman
Andrew Scull discusses how clinical psychology and cognitive behavioral therapy (CBT) emerged after WWII as alternatives to psychoanalysis, becoming dominant through better empirical evidence, shorter treatment duration, and federal funding advantages. CBT's symptom-focused approach proved more measurable and reproducible than psychoanalytic theory, though its effectiveness remains limited for serious mental disorders.
Do anti-depressant drugs work? | Andrew Scull and Lex Fridman
Andrew Scull discusses the history and efficacy of antidepressants, particularly SSRIs like Prozac, explaining that while they statistically outperform placebo, the clinical improvement is often marginal and comes with significant side effects including sexual dysfunction and withdrawal difficulties. He also addresses 'diagnostic creep' in psychiatry, where conditions become increasingly broadly defined and diagnosed over time.