Running evals on your internal agents is how you keep them trustworthy over time
The evaluation of internal AI agents is crucial for maintaining their reliability. An internal eval platform allows engineers to review agent performance regularly and ensure their responses are accurate, particularly in critical areas like coding.
Summary
In the discussion, it is emphasized that regular evaluations of internal AI agents are essential for maintaining trustworthiness over time. Each evaluation generates a log that is handled by engineers who analyze whether the agent's responses are correct and satisfactory. This process mirrors the evaluations conducted for customer-facing AI products, highlighting the importance of maintaining high standards for AI, especially those involved with important tasks such as coding. Continuous assessment allows for ongoing improvements in the performance of these internal tools.
Key Insights
- Regular evaluations of internal agents are logged into an internal eval platform for review by engineers.
- Engineers assess whether agents get answers right or wrong as part of the evaluation process.
- The scoring mechanism of agents is subject to scrutiny during evaluations to ensure quality.
- Internal evaluations mirror those used for customer-facing AI products, emphasizing their importance.
- Special attention is needed for internal AI bots that are involved in critical tasks such as coding.
Topics
Transcript
[0:00] A lot of great folks run evals on this internal agent. So every time this review is run, it gets logged into an internal eval platform and an engineer looks at it and says, "Did the agent get this right? Did the agent get this wrong? Are we happy with the scoring mechanism?" So, very similar to how you'd use evals to improve your customerf facing AI products, you're going to want to use evals to improve your internally facing AI bots, especially ones that touch really critical things [music] like code.
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
Barbie Bench: AGI Has Not Arrived
A content creator demonstrates Claude Opus's Barbie fashion designer 3D rendering project, highlighting both impressive and flawed outputs. While the model successfully created an interactive 3D game where users can dress Barbie and visit a fitting room, it struggled significantly with rendering accurate hands, facial features, and body proportions, leading the creator to conclude that AGI has not yet arrived.
I reviewed Opus 5.5 and GPT-6 Sol live - and the results surprised me
A live review comparing three newly released AI models—Opus 5.5, GPT-6 Soul, and GPT-6 Luna—where the reviewer conducts blind testing across multiple task categories and finds that while Opus 5.5 significantly improves on previous versions, OpenAI's models excel in creative tasks like SVG generation.
Warp’s Productivity dashboard is an eng manager’s dream
Warp's new productivity dashboard centralizes engineering team visibility for managers and CTOs, enabling real-time monitoring of development velocity, code quality improvements, and cost efficiency across distributed teams rather than relying on individual local setups.
She uses Claude Code to find a house with the right vibe
Hillary shares an AI workflow she uses for house hunting that leverages Claude Code to find listings and filter them against her subjective "vibe-specific criteria," which she describes as evaluating whether a space would feel good to wake up in and spend time at.
She hasn’t sent a birthday text herself in 6 months
A speaker shares how they've delegated birthday acknowledgments to AI for the past 6 months, having their AI assistant research friends' children's ages and purchase appropriate gifts on Amazon. They frame this as automating 'emotional labor' through AI technology.