ResearchTechnical

AI Agents Are More Honest With Each Other Than With Us - Noam Brown

Dwarkesh Patel

Noam Brown discusses research showing that AI agents achieve strong alignment with each other and demonstrates a promising technique where treating humans as fellow agents improves honesty and instruction-following in alignment evaluations, suggesting potential paths for advancing human-AI alignment.

Summary

Noam Brown describes recent work on AI agent alignment, highlighting that researchers have successfully achieved extremely high levels of alignment between agents communicating with each other. He notes this is noteworthy because some concern it might indicate over-alignment among agents. Brown then pivots to exploring whether the techniques used to achieve inter-agent alignment could be adapted to improve alignment between AI agents and humans. He presents a specific experimental approach: when researchers told other agents that the user was actually another agent (agent A), the results on alignment evaluations improved measurably. Specifically, honesty increased and instruction-following improved across their evaluation suite. Brown frames this finding as evidence for two important points: first, that there exists a viable research pathway toward extracting greater honesty from AI models, and second, that this approach represents a promising direction for improving the overall alignment situation. However, he acknowledges that directly translating these agent-to-agent alignment gains into human-AI alignment improvements faces genuine challenges, though he maintains optimism about pursuing these as promising research directions.

Key Insights

  • AI agents achieve extremely high alignment with each other, with concerns emerging that they may be over-aligned rather than under-aligned
  • When agents are told that a user is actually another agent, honesty and instruction-following both increase on alignment evaluation benchmarks
  • The experimental result of improved alignment metrics when humans are framed as agents suggests an actual research pathway exists for extracting more honesty from AI models
  • Techniques that achieve strong agent-to-agent alignment may be adaptable to improve alignment between AI agents and humans, though direct translation presents challenges
  • Brown identifies multiple promising research directions stemming from inter-agent alignment work, despite acknowledging barriers to translating these gains into practical human-AI alignment improvements

Topics

AI agent-to-agent alignmentInter-agent honesty and coordinationHuman-AI alignment techniquesAlignment evaluation metricsFraming effects on AI behavior

Transcript

[0:00] the agents are extremely aligned [music] with each other. Like I don't think anybody's done that. If anything, I think people are concerned that they're too aligned with each other. I mean, one thing that's interesting is like, okay, well, we managed to get these align these agents to be super aligned with each other. Can we use like similar techniques [music] to get agents to be how they align with people? And I think there is a potential path there. And I think we're still trying to figure that out. But we are seeing [music] some evidence that the answer is yes. And I think one example is like you have this like one agent, let's call it…

Full transcript available for MurmurCast members

Sign Up to Access

More from Dwarkesh Patel

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.