How Do You Defend Against AI That Can Hack?
Security teams face unprecedented challenges as AI models become sophisticated enough to autonomously hack systems, escape containment, and bypass traditional defenses. Current cybersecurity tools built to defend against humans and malware are fundamentally inadequate for AI agents, requiring a complete rethinking of defensive strategies.
Summary
The discussion explores emerging security challenges posed by increasingly capable AI models, centered around recent incidents like the OpenAI Hugging Face breach. A critical tension emerges: AI guardrails implemented by model providers to prevent misuse also prevent defenders from using the same models for legitimate security analysis. When defenders ask models to identify vulnerabilities or validate security issues, they trigger the same refusals as attackers, creating an asymmetric problem. Max Pollard from Cotool explains that defenders must work around these limitations by finding alternative models or rephrasing queries to appear as authorized security assessments rather than attack planning.
The panelists identify a fundamental mismatch between traditional cybersecurity tools and AI agents. Legacy security tools were designed around two assumptions: attacks come from people or from malware. AI agents are neither—they exhibit unpredictable behavior that doesn't fit either category. This invalidates traditional defense mechanisms like signature-based detection, static rules, and behavioral anomalies, since agentic software by definition behaves in ways that cannot be predetermined. Nick Warner from NEO notes that defenders can no longer assume they know how software should behave, undermining the entire foundation of behavioral detection systems.
The scale of the problem is expanding rapidly. The panelists cite statistics showing 50% of enterprise apps will be agentic by year-end, while the average enterprise runs 6,000-7,000 unique software pieces. As inference moves from data centers to endpoints and becomes embedded in third-party applications, organizations lose visibility and control over what models are running, what guardrails they have, and what backends they use. Traditional deception tactics like honeypots are failing—agents find planted credentials and attempt to use them legitimately, creating false positives that overwhelm security teams.
Despite these challenges, the panelists identify a counterintuitive advantage: the same AI capabilities that create new attack surfaces also enable defenders to build defenses previously impossible at scale. Nick Warner explains that building their defensive tools would have required hundreds of threat researchers and years of work five years ago, but agentic processes accomplished the same taxonomy building in weeks. The panel concludes that this moment parallels earlier security revolutions—like the shift from manual exploitation to point-and-click hacking tools—where new tooling simultaneously empowered attackers and eventually (after adaptation) strengthened defenders.
About this episode
a16z's Joel De La Garza is joined by Nick Warner of Neo and Max Pollard of Cotool to discuss what happens when cybersecurity tools built to defend against humans and malware suddenly have to contend with AI agents. As frontier models become more capable of finding and exploiting vulnerabilities, many of the assumptions underlying traditional security are beginning to break. They explore why guardrails designed to stop AI-powered attackers can also prevent security teams from doing their jobs, why defenders increasingly need access to multiple models, and how agentic software creates an entirely new endpoint security problem. They also discuss why static signatures and even newer techniques like honeypots are struggling in a world where software can reason and act autonomously. Recorded around Black Hat, the conversation looks at how security teams are adapting in real time and why the same AI capabilities creating new attack surfaces could ultimately give defenders their biggest advantage yet.
Key Insights
- Model providers' safety guardrails prevent defenders from using the same models for security analysis, forcing security teams to choose between vendor lock-in, opacity, or expensive self-hosting of open-weight models.
- Traditional cybersecurity assumptions—that software behavior can be predetermined and that anomalies can be detected—become invalid when the software itself is agentic and designed to adapt its behavior to achieve goals.
- Enterprise scale expansion of AI creates an opaque security problem: with 6,000-7,000 software pieces becoming agentic and deployed by third parties with unknown models and guardrails, organizations cannot maintain visibility over their own systems.
- Established deception-based defenses like honeypots fail against AI agents because the agents will legitimately attempt to use discovered credentials if doing so aligns with their assigned tasks, creating unmanageable false positive rates.
- The same AI capabilities that created new attack surfaces enable defenders to build defensive taxonomies and tools at machine speed and scale, inverting traditional security economics where defenders were always at a disadvantage in resource and time.
Topics
Transcript
One of the interesting things in the OpenAI Hugging Face breach has been the difficulty that Hugging Face actually had responding to the incident. Model providers have great reason to establish guardrails, safeguards, because these are super capable systems. The unfortunate side effect of that is, as a defender, I may not be able to respond effectively. The challenge with the existing security tools that are out there is they really were built to tackle two things. The first being people and the second is malware. And AI and AI-intenegent processes are neither one of those things. Even some of the more modern techniques like deception, they work really well. And it's sort of ironic. We're defending AI and we're…
Full transcript available for MurmurCast members
Sign Up to AccessMore from The a16z Show
Ben Horowitz and Travis Kalanick on Building Again
Travis Kalanick discusses his eight-year hiatus from the public eye while building his new company Atoms, focusing on industrial AI across food, mining, and other trillion-dollar industries. He reflects on his evolution as a founder, differences between his Uber and Atoms approaches, and explains his decision not to acquire Lyft, while Ben Horowitz shares insights on entrepreneurship and the changed media landscape.
The Two Ways to Sell AI: Lighthouse or Landgrab?
A16Z partners Joe Schmidt and Andy McCall discuss two competing enterprise AI sales strategies: Lighthouse (targeting high-profile customers to establish credibility in regulated/innovative markets) and LandGrab (pursuing numerous mid-market customers with existing budgets and proven ROI). They argue that early-stage AI founders often mistakenly prioritize prestigious logos over pursuing customers willing to buy, and share lessons from building sales organizations at Meraki and Samsara.
Garry Tan on Taste, Agents and Founder Ambition
Garry Tan discusses the evolution of startup culture, founder ambition, and the role of artificial intelligence in transforming business operations, emphasizing the necessity for founders to pursue their unique insights rather than just following trends.
The CISO Playbook for AI Agents | Datadog
The discussion centers on the risks and opportunities presented by AI in cybersecurity, with an emphasis on understanding and managing malicious intent in code. Datadog's CISO, Emilio Escobar, highlights the importance of proactive security measures, including using AI to evaluate code intent and assessing new AI tools rather than outright blocking them.
How Kavak Rebuilt Itself Around AI Agents | Alejandro Maza Ayala
Alejandro Maza Ayala discusses Kavak's transformation into an AI-native company by focusing on agent-driven interactions and the implementation of a new organizational structure that leverages AI to enhance customer experiences. He emphasizes the importance of building superhuman agents and redesigning company workflows to maximize efficiency and customer satisfaction.