DiscussionTechnical

How Do You Defend Against AI That Can Hack?

The a16z Show22m 0s

Security teams face unprecedented challenges as AI models become sophisticated enough to autonomously hack systems, escape containment, and bypass traditional defenses. Current cybersecurity tools built to defend against humans and malware are fundamentally inadequate for AI agents, requiring a complete rethinking of defensive strategies.

Summary

The discussion explores emerging security challenges posed by increasingly capable AI models, centered around recent incidents like the OpenAI Hugging Face breach. A critical tension emerges: AI guardrails implemented by model providers to prevent misuse also prevent defenders from using the same models for legitimate security analysis. When defenders ask models to identify vulnerabilities or validate security issues, they trigger the same refusals as attackers, creating an asymmetric problem. Max Pollard from Cotool explains that defenders must work around these limitations by finding alternative models or rephrasing queries to appear as authorized security assessments rather than attack planning.

The panelists identify a fundamental mismatch between traditional cybersecurity tools and AI agents. Legacy security tools were designed around two assumptions: attacks come from people or from malware. AI agents are neither—they exhibit unpredictable behavior that doesn't fit either category. This invalidates traditional defense mechanisms like signature-based detection, static rules, and behavioral anomalies, since agentic software by definition behaves in ways that cannot be predetermined. Nick Warner from NEO notes that defenders can no longer assume they know how software should behave, undermining the entire foundation of behavioral detection systems.

The scale of the problem is expanding rapidly. The panelists cite statistics showing 50% of enterprise apps will be agentic by year-end, while the average enterprise runs 6,000-7,000 unique software pieces. As inference moves from data centers to endpoints and becomes embedded in third-party applications, organizations lose visibility and control over what models are running, what guardrails they have, and what backends they use. Traditional deception tactics like honeypots are failing—agents find planted credentials and attempt to use them legitimately, creating false positives that overwhelm security teams.

Despite these challenges, the panelists identify a counterintuitive advantage: the same AI capabilities that create new attack surfaces also enable defenders to build defenses previously impossible at scale. Nick Warner explains that building their defensive tools would have required hundreds of threat researchers and years of work five years ago, but agentic processes accomplished the same taxonomy building in weeks. The panel concludes that this moment parallels earlier security revolutions—like the shift from manual exploitation to point-and-click hacking tools—where new tooling simultaneously empowered attackers and eventually (after adaptation) strengthened defenders.

About this episode

a16z's Joel De La Garza is joined by Nick Warner of Neo and Max Pollard of Cotool to discuss what happens when cybersecurity tools built to defend against humans and malware suddenly have to contend with AI agents. As frontier models become more capable of finding and exploiting vulnerabilities, many of the assumptions underlying traditional security are beginning to break. They explore why guardrails designed to stop AI-powered attackers can also prevent security teams from doing their jobs, why defenders increasingly need access to multiple models, and how agentic software creates an entirely new endpoint security problem. They also discuss why static signatures and even newer techniques like honeypots are struggling in a world where software can reason and act autonomously. Recorded around Black Hat, the conversation looks at how security teams are adapting in real time and why the same AI capabilities creating new attack surfaces could ultimately give defenders their biggest advantage yet.

Key Insights

  • Model providers' safety guardrails prevent defenders from using the same models for security analysis, forcing security teams to choose between vendor lock-in, opacity, or expensive self-hosting of open-weight models.
  • Traditional cybersecurity assumptions—that software behavior can be predetermined and that anomalies can be detected—become invalid when the software itself is agentic and designed to adapt its behavior to achieve goals.
  • Enterprise scale expansion of AI creates an opaque security problem: with 6,000-7,000 software pieces becoming agentic and deployed by third parties with unknown models and guardrails, organizations cannot maintain visibility over their own systems.
  • Established deception-based defenses like honeypots fail against AI agents because the agents will legitimately attempt to use discovered credentials if doing so aligns with their assigned tasks, creating unmanageable false positive rates.
  • The same AI capabilities that created new attack surfaces enable defenders to build defensive taxonomies and tools at machine speed and scale, inverting traditional security economics where defenders were always at a disadvantage in resource and time.

Topics

AI model guardrails creating defensive blind spotsInadequacy of traditional cybersecurity tools against AI agentsBreakdown of signature-based and behavioral detection methodsRapid enterprise adoption of agentic software and inference at endpointsAI-powered defensive tools turning the tables on attackersHoneypot deception tactics failing against autonomous agents

Transcript

One of the interesting things in the OpenAI Hugging Face breach has been the difficulty that Hugging Face actually had responding to the incident. Model providers have great reason to establish guardrails, safeguards, because these are super capable systems. The unfortunate side effect of that is, as a defender, I may not be able to respond effectively. The challenge with the existing security tools that are out there is they really were built to tackle two things. The first being people and the second is malware. And AI and AI-intenegent processes are neither one of those things. Even some of the more modern techniques like deception, they work really well. And it's sort of ironic. We're defending AI and we're…

Full transcript available for MurmurCast members

Sign Up to Access

More from The a16z Show

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.