How to make really good Claude skills (clearly explained in 42 seconds)
The transcript outlines a method for creating high-quality Claude skills by providing context, defining clear criteria, iterating until successful, and letting the AI write the skill itself. The process emphasizes testing and collaborative debugging to ensure reliability. The speaker argues AI-written skills outperform human-written ones because they reflect what actually worked.
Summary
The speaker addresses a common problem with Claude skills: average output quality resulting from a lack of personalized context. The solution begins with grounding the skill in specific, real-world criteria. Using a lead research agent as an example, the speaker suggests instructing the AI to check sources like Twitter, YouTube, and Trustpilot, and to reject leads immediately if two or more sources are missing or unfavorable. This clarity of definition sets up the next phase effectively.
The second major phase is iteration. The speaker recommends running the process multiple times until a clean, end-to-end run is achieved. Once that happens, the key insight is to have the AI itself review what it just did and convert that successful run into a formalized skill. The speaker's argument is that the AI will write a better skill than a human would, precisely because it has firsthand knowledge of what actually worked during the iterations.
Finally, the speaker emphasizes rigorous testing. If the skill breaks, the recommendation is to ask the AI why it failed, fix the issue collaboratively, and test again. This iterative debugging loop is framed as essential to ensuring the skill remains robust and does not break in the future. The overall framework is: contextualize, iterate, let AI formalize, and test collaboratively.
Key Insights
- The speaker argues that Claude skills produce average output by default because they lack the context of the user's specific style and criteria.
- The speaker uses a lead research agent example, specifying that checking Twitter, YouTube, and Trustpilot — and rejecting leads if two are missing or poor — is how you define acceptable standards within a skill.
- The speaker claims that running multiple iterations until a clean end-to-end run is achieved is a necessary step before formalizing any skill.
- The speaker argues that having the AI review its own successful run and write the skill from that experience produces better results than a human writing the skill manually.
- The speaker frames collaborative debugging — asking the AI why something broke, fixing it together, and retesting — as the critical step to ensuring a skill never breaks again.
Topics
Transcript
[0:00] Every time you create a claude skills, the output is often average. That's because skills lack the context of your style. Say you're building a lead research agent. Tell it to check their Twitter, YouTube, and Trust Pilot. If two are missing or look bad, reject instantly by defining what's okay and what's not. You set it up for the next step, iteration. Run multiple iterations until you get one clean run end to end. Then ask the AI to review what it just did and turn that into a skill. AI writes a skill better than you will [0:30] because it now knows what actually worked. Finally, test the skill. If something breaks, ask it why. Fix it…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Greg Isenberg
Muse AI Connectors: The Next App Store Moment?
Greg Eisenberg analyzes Meta's new Muse AI connectors as a potential "app store moment" for AI, drawing parallels to Apple's 2008 App Store launch. He provides four startup ideas, explains how to build connectors using coding agents, and outlines distribution strategies beyond relying on Meta's directory featuring.
$30M Writer: Never write AI slop again
Nicholas Cole, who earned $30M+ writing online, explains how to create valuable content in the AI era by building a personal language model of your unique ideas, personality, and perspective rather than generic commodity content. He argues that voice, original thinking, and consistent association through personality details create sustainable competitive advantage, while AI should be used to repeat and remix existing ideas, not generate net new thinking.
Jev is HERE. How to use it
Jev is a new AI classifier model created by Dooo Almeida that makes fast, low-cost decisions on structured inputs rather than generating text like traditional LLMs. It processes information through a defined schema to return probability-based classifications across multiple categories, enabling applications from email triage to content clipping to flight booking automation.
Instinct AI is For Real. What You Need to Know.
Hosts Greg and Remy review Instinct AI, an invite-only personal agent app raising at a $10 billion valuation, demonstrating its capabilities for booking flights, restaurants, obtaining visas, and other personal tasks. They highlight its impressive simplicity, human-like communication style, and network effects through its trusted agent connectivity, while flagging privacy and data retention concerns.
Building a Software Factory that actually works (Full Course)
Ross Mickey explains how to build a software factory—a systematic workflow using AI agents to ship high-quality software at scale. The factory concept uses markdown files to define agent behavior and follows four key steps: isolate (branching), build (with code structure guidelines), prove (with before/after testing), and ship (with automated code review).