TechnicalDiscussion

Yann LeCun: Deep Learning, ConvNets, and Self-Supervised Learning | Lex Fridman Podcast #36

Lex Fridman

Yann LeCun discusses deep learning's fundamentals, the importance of self-supervised learning for building world models, and the necessary components for achieving human-level AI. He argues that current AI systems lack true understanding and that future progress requires machines to learn predictive models of the world similar to how infants learn through observation.

Summary

In this wide-ranging conversation, Yann LeCun covers multiple interconnected themes in AI development. He begins by discussing HAL 9000 from 2001: A Space Odyssey, framing the AI alignment problem as a value misalignment issue similar to designing legal systems for humans—we must constrain objective functions through rules and oversight. LeCun emphasizes this is not a new problem but a continuation of millennia-old practices of lawmaking.

On the history of deep learning, LeCun explains why neural networks fell out of favor in the 1990s. Key obstacles included: difficulty implementing backpropagation without modern languages like Python, lack of software infrastructure for experimentation, insufficient knowledge of training tricks, and the inability to share code due to pre-open-source licensing practices. He describes how his team at Bell Labs had to write a Lisp interpreter with a compiler to C to build production character recognition systems, an enormous undertaking that deterred others from pursuing neural networks.

LeCun challenges the notion of "general" intelligence, using a thought experiment about rewiring visual input to demonstrate that human intelligence is heavily specialized to process locality in the natural world. He argues we have roughly 2^(2^1,000,000) possible boolean functions our visual cortex could theoretically compute, but can only implement a tiny fraction due to architectural constraints. Human intelligence appears general only because it's general to all things we can conceive of—a much smaller space than all possible things.

On unsupervised learning, LeCun reframes the concept as self-supervised learning, where machines predict masked or hidden portions of input rather than learning without any supervision signal. He describes successful applications in NLP (like BERT-style systems that predict missing words) and discusses why this approach is harder for images and video—because there are multiple valid completions, making uncertainty representation challenging. He explicitly rejects active learning as transformative, suggesting it will only incrementally improve existing approaches.

A central thesis throughout is the necessity of learning predictive models of the world. LeCun argues this is crucial for intelligent autonomous systems and notes that even infants learn basic physics concepts (gravity, object stability) through observation and light interaction. He contrasts this with current reinforcement learning approaches, which require enormous amounts of trial and error (e.g., 80 hours of training to match 15 minutes of human learning in Atari games, or 200 years of simulated play for AlphaStar). Model-based learning with world prediction would allow systems to avoid catastrophic errors like driving off cliffs by using internal simulation.

Regarding autonomous vehicles, LeCun acknowledges deep learning will be fundamental but notes current approaches heavily engineer the problem (mapping, specialized sensors, constrained geography). Long-term solutions will rely increasingly on learned models, but this remains unsolved territory.

On the path to human-level AI, LeCun identifies multiple obstacles but admits uncertainty about how many remain. He identifies self-supervised world modeling as the first major peak to climb. The architecture he envisions for intelligent autonomous systems requires four components: (1) a predictive model of the world, (2) an objective function predictor (analogous to the basal ganglia), (3) a policy network that determines optimal actions, and (4) a hard-wired contentment objective. He emphasizes emotions are not luxuries but essential to intelligence, as they represent drives and contentment predictions that shape behavior.

LeCun critiques over-claiming in AI, particularly regarding projects like Sophia the robot, which he views as misleading the public about capabilities. However, he also notes that most academic peer review and published research is reasonably careful about unsupported claims. He advocates for grounding AI systems in the world through visual perception, text, video, and potentially virtual interaction—not necessarily requiring physical embodiment but requiring some understanding of how the world works to avoid frustrating users.

Key Insights

  • LeCun argues that AI alignment is fundamentally similar to legal system design—both involve constraining objective functions through rules to prevent harmful outcomes, and this practice extends back millennia.
  • The decline of neural networks in the 1990s resulted not from theoretical flaws but from practical engineering barriers: lack of accessible programming tools, inability to share code due to IP restrictions, and absence of training knowledge that only deep learning researchers possessed.
  • Human intelligence is not generally intelligent but highly specialized—our visual cortex can only compute a vanishingly small fraction of theoretically possible functions due to architectural constraints optimizing for local spatial relationships in the natural world.
  • Self-supervised learning succeeds in NLP partly because predicting discrete words from limited vocabularies allows straightforward uncertainty representation, but fails for images and video because multiple valid reconstructions require representing continuous high-dimensional distributions.
  • Intelligent autonomous systems require four components: a world predictive model, an objective function predictor, a policy network, and a hard-wired contentment objective—emotions are not optional features but essential to directing intelligent behavior.

Topics

Self-supervised learning and world modelsValue alignment and objective function designHistory of deep learning and neural network skepticismSpecialization vs. generality in intelligenceModel-based reinforcement learning vs. model-free approachesPredictive uncertainty representation in images and videoAutonomous systems architectureGrounding and common sense reasoningEmotions and contentment prediction in AI

Transcript

[0:00] the following is a conversation with Jana kun he's considered to be one of the fathers of deep learning which if you've been hiding under a rock is the recent revolution in AI that's captivated the world with the possibility of what machines can learn from data he's a professor in New York University a vice president and chief AI scientist a Facebook & Co recipient of the Turing Award for his work on deep learning he's probably best known as the founding father of convolutional neural networks in particular their application [0:32] to optical character recognition and the famed M NIST data set he is also an outspoken personality unafraid to speak his mind in a distinctive French…

Full transcript available for MurmurCast members

Sign Up to Access

More from Lex Fridman

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.