OpinionTechnical

The Problem With How LLMs Generate Text - Grant Sanderson

Dwarkesh Patel

Grant Sanderson explains that autoregressive text generation in LLMs is fundamentally misaligned with human writing processes. Rather than composing with foresight and deep connections, LLMs are forced to predict tokens sequentially with memory wiped between predictions, making it difficult for them to form the unlikely but substantive connections that characterize good writing.

Summary

Sanderson uses an analogy to illustrate the oddness of autoregressive generation: imagine an intelligent person locked in a box who can only interact by receiving slips of paper asking them to predict what comes next, then having their memory wiped repeatedly. The resulting essay would be poor despite the person's intelligence, because sequential prediction is fundamentally different from compositional thinking. He argues that LLMs become slaves to immediate context—drawing connections within a narrow field—but the truly valuable connections are inherently unlikely and unpredictable as next tokens. While reinforcement learning can improve performance in some ways, there's no specific incentive mechanism driving the model toward making unlikely but substantive connections when the vast majority of unlikely connections aren't the statistically predictable next token. Sanderson concludes that this suggests intelligent capability may be locked inside the model, but autoregression is simply a weird constraint on how that intelligence can interact with and produce text.

Key Insights

  • Autoregressive generation fundamentally differs from human compositional writing because humans plan and think through ideas holistically, while LLMs predict tokens sequentially with no memory between steps
  • LLMs become enslaved to immediate context when answering questions, drawing connections narrowly within a particular field rather than making the substantive cross-domain connections that characterize good writing
  • The most valuable and substantive connections in writing are inherently unlikely as next-token predictions, creating a misalignment between what makes good writing and what autoregressive models are optimized to predict
  • Reinforcement learning improvements to LLMs lack specific incentive mechanisms that would push the model toward generating unlikely-but-substantive connections, since these connections contradict the statistical likelihood objective
  • There may be significant intelligence locked inside LLMs, but the autoregressive interaction mechanism is an inherently limiting constraint on how that intelligence can be expressed through text generation

Topics

Autoregressive text generation limitationsMismatch between LLM generation and human compositionContext dependence and token predictionUnlikely connections and substantive thinkingReinforcement learning and incentive structures

Transcript

[0:00] Auto regression is actually like a really really weird [music] way to produce stuff if you think about it. Like you're an intelligent person. Imagine I locked you in a box and then the only way that you have of interacting with the world is that you receive a slip of paper and then someone says, "Can you like predict what will come next?" And then you predict what will come next and then your memor is [music] wiped. Imagine that was done a whole bunch and then what comes out on the other end? They're like, "Look at this essay that [music] you wrote." You might look at that and be like, "This is awful. That's not the…

Full transcript available for MurmurCast members

Sign Up to Access

More from Dwarkesh Patel

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.