Running evals on your internal agents is how you keep them trustworthy over time
The evaluation of internal AI agents is crucial for maintaining their reliability. An internal eval platform allows engineers to review agent performance regularly and ensure their responses are accurate, particularly in critical areas like coding.
Summary
In the discussion, it is emphasized that regular evaluations of internal AI agents are essential for maintaining trustworthiness over time. Each evaluation generates a log that is handled by engineers who analyze whether the agent's responses are correct and satisfactory. This process mirrors the evaluations conducted for customer-facing AI products, highlighting the importance of maintaining high standards for AI, especially those involved with important tasks such as coding. Continuous assessment allows for ongoing improvements in the performance of these internal tools.
Key Insights
- Regular evaluations of internal agents are logged into an internal eval platform for review by engineers.
- Engineers assess whether agents get answers right or wrong as part of the evaluation process.
- The scoring mechanism of agents is subject to scrutiny during evaluations to ensure quality.
- Internal evaluations mirror those used for customer-facing AI products, emphasizing their importance.
- Special attention is needed for internal AI bots that are involved in critical tasks such as coding.
Topics
Transcript
[0:00] A lot of great folks run evals on this internal agent. So every time this review is run, it gets logged into an internal eval platform and an engineer looks at it and says, "Did the agent get this right? Did the agent get this wrong? Are we happy with the scoring mechanism?" So, very similar to how you'd use evals to improve your customerf facing AI products, you're going to want to use evals to improve your internally facing AI bots, especially ones that touch really critical things [music] like code.
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
Enterprise AI fails on governance, not the model
The speaker discusses the importance of skill management and governance in enterprise AI platforms, emphasizing that enabling people to build skills is insufficient without proper monitoring, maintenance, and telemetry. The focus is on providing automatic platform-driven suggestions to keep skills current and using data insights to promote high-value skills while deprecating low-value ones.
Onboard your employees to Cowork in 15 minutes with this system
The transcript describes a workstation operating system that streamlines employee onboarding through an intuitive chat-based interface. The system automates tool connections, role confirmation, colleague mapping, calendar integration, and personalization to get new employees productive from day one.
Your prompt is a spec, and a good spec is the whole game
A prompt acts as a specification that defines quality criteria for AI outputs. By clearly articulating 5-10 defining characteristics of what makes something good (whether a garment, photo, or illustration), you provide the AI with the same level of detail a human creator would need to produce excellent results.
How I use ChatGPT to run my fashion business
Yana Welander demonstrates how she built an AI-native fashion brand using ChatGPT and Codex as her technical co-founder, automating design, production, and business operations from sketches to manufacturing and e-commerce. She shares how AI unlocks previously impossible garments while highlighting remaining challenges in pattern generation and realistic image rendering.
The biggest barrier to AI adoption is not fear of technology. It's the absence of muscle memory.
The speaker argues that the primary barrier to AI adoption is not fear but lack of muscle memory—the habitual use of AI tools. Success requires building collaborative practices through empathy, accessible training, and establishing regular workflows rather than overcoming technological anxiety.