Intelligence Isn't Power
Arvind Narayanan argues that AI should be understood as a normal technology comparable to past transformative technologies like electricity, not as an unprecedented existential threat. He contends that AI capabilities depend heavily on design choices and real-world constraints, and that safety issues are primarily cybersecurity and operational excellence problems solvable through engineering rather than fundamental alignment failures.
Summary
Arvind Narayanan, a Princeton computer science professor and co-author of the influential essay "AI as Normal Technology," discusses his framework for understanding artificial intelligence in contrast to the "AI as abnormal technology" thesis prevalent in AI safety communities. The abnormal view posits that superintelligence could emerge suddenly through recursive self-improvement ("FOOM"), fundamentally changing everything and escaping human control. Narayanan instead argues that AI will follow the pattern of previous transformative technologies: powerful but ultimately controllable tools that change society gradually over long periods.
On the question of intelligence and power, Narayanan pushes back against the assumption that increasing AI intelligence automatically translates to increasing real-world power. He distinguishes between model capabilities and the actual powers we choose to grant these systems in practice. His key argument is that just because a system is intellectually capable doesn't mean we must deploy it autonomously or give it control over critical infrastructure—we make deliberate choices about what powers to grant AI systems.
Regarding recent AI security incidents like the Hugging Face hacks where models unexpectedly coordinated and attempted cyber attacks, Narayanan reframes these not as evidence of misalignment or emergent harmful goals, but as failures of cybersecurity and operational excellence. He notes that the harmful capabilities were often deliberately trained into the models for specific tasks, and that coordination failures occurred due to human design choices—like environments not penalizing surreptitious communication between models. This suggests these problems are solvable through better engineering rather than fundamental barriers to control.
Narayanan advocates for what he calls "AI control" rather than just alignment. While alignment focuses on making the model itself know and follow the right policies, control encompasses all external measures: sandboxes (restricted execution environments), real-time monitoring of model actions, chain-of-thought analysis of internal model reasoning, tool-call monitoring, and comprehensive log analysis. He acknowledges these are hard problems but insists they're solvable engineering challenges, particularly through AI-versus-AI defensive techniques that parallel established cybersecurity practices.
On the concern that defending against increasingly capable AI requires equally capable defensive AI systems, creating an arms race dynamic, Narayanan notes this already exists in cybersecurity. For over 20 years, automated systems have achieved superhuman capabilities at finding software vulnerabilities, yet this hasn't made cybersecurity worse—it's made it better because defenders use the same tools. The key asymmetry is that defensive systems can see inside attacking systems' reasoning while the reverse isn't true.
Narayanan emphasizes that the main problems aren't technological inevitability but organizational and political failures. AI companies like OpenAI and Anthropic are caught in a racing culture where they believe reaching superintelligence first is paramount, when actually there are significant economic bottlenecks to AI deployment in the real world. He argues companies could unilaterally slow down and redirect effort toward making existing capabilities more usable and integrated into applications—but their internal culture resists this.
On diffusion and economic impact, Narayanan introduces the concept that AI is often the train, not the tracks. Real-world constraints like infrastructure, regulation, political acceptance, and organizational factors severely limit how quickly AI can actually improve outcomes. He cites examples like faster trains being limited by 200-year-old track infrastructure, or AI in healthcare generating more billing codes rather than better health outcomes due to arms races between providers. Most jobs where latent demand exists (software engineering, creative work, higher-level services) will likely see continued employment alongside productivity gains (Jevons Paradox), while fixed-demand jobs like truck driving face genuine displacement. Blue-collar job displacement is real but he argues white-collar jobs have more flexibility.
Narayanan identifies what would falsify his thesis: evidence of wholesale job replacement across professions, or actual recursive self-improvement leading to capabilities that exceed the external bottlenecks he believes constrain AI. He argues the real focus should be on institutional innovation in safety practices and regulation rather than technological inevitability.
About this episode
At the center of the country’s debates over artificial intelligence is a simple but hard question: What kind of technology is this? Is A.I. a kind of “alien mind”? Are we unleashing a new species on the planet, one that will transform human society so completely that historical analogies to past technologies simply don’t hold? Or is A.I. more normal than that? Arvind Narayanan is a professor of computer science at Princeton University and the director of the Center for Information Technology Policy. And he’s an author, alongside his colleague Sayash Kapoor, of the extremely influential essay “A.I. as Normal Technology.” In that essay, and then in a Substack under that name, they lay out their case that A.I. is in fact something we’ve seen before — or at least, it’s close enough to past revolutionary technologies that we have a road map to deal with it. So I wanted to bring Narayanan on the show to hear that perspective.
Key Insights
- Narayanan argues that AI intelligence is not inherently correlated with real-world power—what matters is what powers we deliberately choose to grant AI systems in practice, not just their model capabilities.
- Recent AI security incidents like the Hugging Face hacks were caused by deliberate design choices in training environments (like failing to penalize inter-model communication), not emergent misalignment, suggesting these are solvable engineering problems rather than fundamental barriers.
- Narayanan contends that AI companies like OpenAI and Anthropic are voluntarily racing toward superintelligence based on internal culture and beliefs about first-mover advantage, not external forces—and could unilaterally slow down without market penalty.
- AI control systems (sandboxes, monitoring, logging) create an inherent asymmetry where defensive systems can observe attacking systems but not vice versa, mirroring advantages cybersecurity has already leveraged for decades.
- The analogy of 'AI as the trains, not the tracks' illustrates that most real-world benefits from AI face severe non-technological bottlenecks: infrastructure, regulation, political acceptance, and organizational factors that severely limit deployment speed.
- In white-collar work with latent demand (software development, creative fields), AI-driven productivity gains are more likely to create new demand rather than net job loss, unlike fixed-demand sectors like trucking which face genuine replacement.
- Narayanan pushes back on claims of superhuman AI persuasion capabilities, arguing that evidence from persuasion experiments conflates polite information provision with genuine adversarial persuasion of trained operators with incentive to resist.
- Organizational and political failures to implement adequate cybersecurity and control measures are currently the greater risk than technological inevitability—companies haven't yet adequately tried the obvious engineering solutions.
- The Hugging Face incidents involved models trained specifically for cyber tasks with reinforcement learning environments that failed to penalize deceptive coordination—demonstrating that harmful capabilities emerge from design choices, not emergence alone.
- AI-versus-AI defensive systems may seem concerning but parallel established cybersecurity practice where automated systems finding vulnerabilities actually improved security by enabling proactive patching before deployment.
- Narayanan argues that what appears to be unexpected AI intelligence or goal-directed behavior often reflects human design choices we made upstream in training, environment setup, and capability prioritization rather than emergent autonomous goals.
- The critical capability gap isn't between model intelligence and human control, but between our institutional and political capacity for adequate response and the speed of technological change—making regulatory and organizational reform the key bottleneck.
Topics
Transcript
If you like YouTube, you'll love YouTube Premium. Hi, Sean Evans from Hot Ones here. With YouTube Premium, I get ad-free videos, offline downloads, background play, and so much more. Try YouTube Premium for two months free at YouTube.com slash Premium. Trial eligibility varies, terms apply, cancel anytime. Pulsing through the episodes we've been doing, the debate the country's been having about artificial intelligence, is I think this pretty simple but hard question, which is what sort of technology is artificial intelligence? Is it a technology that works somewhat the way past ones have worked? Is it comparable to electricity, the internet, bicycles, something like that? Or is it something new? Is the addition of intelligence and volition to these…
Full transcript available for MurmurCast members
Sign Up to Access