How to build a custom AI harness with Claude SDK
A custom AI harness is code wrapped around an AI agent to make it more effective for specific workflows. The speaker demonstrates building a Sentry bug-debugging harness using Claude SDK with a terminal UI, showing how structured constraints, custom prompts, and specific tool adapters enable agents to handle complex tasks more efficiently than general-purpose AI tools.
Summary
The episode demystifies the concept of an AI harness, explaining it as simply code around an AI agent designed to make it more effective for specific use cases. A harness consists of three core components: specific context, the ability to take specific actions, and goals for specific outcomes. The speaker built a practical example: a Sentry bug-debugging harness that investigates production issues, identifies root causes, and generates structured reports.
The harness is appropriate when workflows are repetitive, require specific setup, and have consistent outcomes—such as fixing bugs, managing production incidents, handling support escalations, or other structured processes. The speaker chose Sentry debugging because it involves coding, requires custom context, and has specific desired outcomes like creating Linear tickets and writing follow-up documentation.
The implementation uses Claude SDK as the core agent framework, connected to real tools including Sentry, Vercel, Linear, and GitHub. Rather than using general-purpose tools, the harness includes opinionated adapters that make specific API calls tailored to bug investigation rather than exposing all functionality. The interface is a terminal UI built with the Ink library, allowing intuitive navigation through investigation runs, evidence gathering, and artifacts.
Key architectural features include: custom prompts that encode the specific workflow rather than relying on natural language requests each time, tool policies that restrict what the agent can execute, structured artifact outputs that document findings in a consistent format, and the ability to run via both interactive UI and command-line tools. The speaker emphasized that building the harness required very specific prompting when using general models like GPT and Opus, as they initially resisted incorporating AI and wanted to build purely deterministic systems.
The artifact store saves evidence from each run to the file system, enabling the agent to learn from previous investigations. Output includes task runs with all messages, investigation briefs, relevant logs, worker actions, and summaries—plus formatted HTML reports showing what happened during investigation. The speaker demonstrates a real run investigating dropped edit operations, where the harness identified the issue, ranked potential root causes, and recommended whether to create a Linear issue or attempt a fix.
Key Insights
- A harness is simply code around an AI agent that makes it more effective, and the mystery around the term has been overblown—it's fundamentally about writing code to wrap AI calls to make them more efficient at specific jobs rather than open-ended tasks
- Building a custom harness for specific workflows produces better results than using general-purpose AI tools because the harness removes the need to re-explain intent and allows encoding of precise step-by-step flows that repeat consistently over time
- When building a harness with general AI models like GPT and Opus, they initially resisted incorporating AI and wanted to build purely deterministic systems, requiring very specific prompting to achieve the desired agentic behavior
- Opinionated tool adapters are more effective than exposing general MCPs—rather than letting the agent wander through all available options, the speaker built specific connectors that precisely define what data to pull from APIs in the context of bug investigation
- The speaker realized that by constraining AI work through specific harnesses and workflows, agents can solve very specific problems efficiently, changing the paradigm from open chat fields to structured, orchestrated agent work with custom outcomes
Topics
Transcript
[0:00] A harness is some code around an AI agent that makes it more effective. Why we've seen people build these specific use case harnesses is sometimes with a specific job, you just want to micromanage a little bit. You just want to be more prescriptive about how that job gets done. [music] I'm going to show you how it works and then we will talk about how I built it. So the interface I built for my harness is a terminal UI. The harness core is run on clawed agent [0:31] SDK and then it's connected to real tools. So it's connected [music] to Sentry Vcel and then it's connected to linear and GitHub in terms of getting [music]…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
AI picks my thumbnails now
A content creator demonstrates using OpenAI's Vision API to automatically analyze video frames and select the most flattering ones for YouTube thumbnails. The frame picker app scores frames on expression, clarity, composition, and impact, then combines selected frames with DALL-E 3 to generate polished thumbnail designs.
Codex built my weekly Spotify playlist from Reddit
A developer created an automated website that solves shared Spotify access issues by building personalized weekly playlists from Reddit's music communities. The system scrapes popular music recommendations from Reddit, ranks them by voting, and automatically generates a fresh playlist every Monday morning.
How the OpenAI team uses ChatGPT Sites daily
Kat, a website product leader at OpenAI, demonstrates how the ChatGPT Sites feature with plugin connectors enables diverse use cases ranging from incident management and team collaboration to creative projects like music discovery and 3D game development. The discussion highlights how Sites serves as both a practical business tool and an infrastructure platform for rapid prototyping and creative expression.
Jev beat an LLM at blitz chess
Jev, an AI system, defeats an LLM at blitz chess by using a two-step analysis process that evaluates the top three moves and their subsequent branches in under a second. While LLMs could eventually solve the same problem, they would require significantly more computational resources and time, making Jev's specialized approach more efficient for time-constrained decision-making tasks.
Jev mapped voice to color over the weekend
Jev built a real-time application that maps voice input to colors and emotions using OpenAI's real-time voice API, Java, and a quotes API. The system analyzes emotional tone and returns corresponding colors and relevant quotes in real time.