Humans are still the bottleneck in Warp’s AI factory
Warp discusses how human code review has become the main bottleneck in their AI-assisted software development process, with a 3.5-hour delay from PR to first human review compared to 35 minutes from launch to PR. They're evolving their workflow to reduce human dependency by allowing requesters to review agent-generated code themselves, and plan to eventually skip review entirely for low-risk tasks by treating code review as a risk management exercise.
Summary
Warp's development team identifies a significant inefficiency in their AI-driven software development workflow. While their AI agent can produce code and create pull requests in just 35 minutes from launch, the process then stalls for 3.5 hours waiting for human review—making humans the primary bottleneck in their development cycle. The team previously followed a two-person review model where Person A would create code using an AI agent and Person B would review it. They've since evolved this workflow so that the person requesting the agent's work can review the generated code themselves, reducing coordination overhead. However, they acknowledge this still doesn't provide full confidence in the process. Looking forward, Warp expects their approach to mature further by gradually building confidence in skipping human code review for certain categories of tasks based on risk level. The team frames this evolution as a transition where code review transforms from a mandatory step into a strategic risk management exercise—selectively applied based on task complexity and risk profile rather than applied uniformly to all pull requests.
Key Insights
- Warp's AI can produce code and create pull requests in 35 minutes, but humans take 3.5 hours to review—making human review the actual bottleneck despite AI acceleration elsewhere
- Warp changed their workflow so the person requesting AI-generated code reviews it themselves rather than requiring a separate reviewer, eliminating coordination overhead
- Despite workflow improvements, Warp still lacks full confidence in skipping human code review entirely for all tasks
- Warp plans to gradually skip code review for certain percentages of tasks as confidence builds, rather than applying human review universally
- Warp expects code review to evolve into a risk management exercise rather than a mandatory process applied to all pull requests
Topics
Transcript
[0:00] People are slowing down the software development process a bit . Can we have a little discussion about whether people are really a " bottleneck"? Because if you look at the time from launch to PR, it's 35 minutes. But if from PR to the first human view, it's 3.5 hours. And it's so funny that the cycle still takes so long, man. This is what is really worth focusing on. So it's on this number. They are still a " bottleneck". We still do human code review. Currently, all of our PRs undergo human code review. Previously, the workflow was like this: [0:30] Person A on our team creates something with an agent, and Person B reviews the…
Full transcript available for MurmurCast members
Sign Up to AccessMore from How I AI
Warp agents open PRs to fix the factory itself
Programming agents can autonomously improve factory systems by analyzing failed launches and proposing specific updates to agent definitions. A self-improvement loop enables observer agents to detect failures and generate evidence-based modifications that prevent recurring issues, such as changing specific steps in factory agent procedures.
Barbie Bench: AGI Has Not Arrived
A content creator demonstrates Claude Opus's Barbie fashion designer 3D rendering project, highlighting both impressive and flawed outputs. While the model successfully created an interactive 3D game where users can dress Barbie and visit a fitting room, it struggled significantly with rendering accurate hands, facial features, and body proportions, leading the creator to conclude that AGI has not yet arrived.
I reviewed Opus 5.5 and GPT-6 Sol live - and the results surprised me
A live review comparing three newly released AI models—Opus 5.5, GPT-6 Soul, and GPT-6 Luna—where the reviewer conducts blind testing across multiple task categories and finds that while Opus 5.5 significantly improves on previous versions, OpenAI's models excel in creative tasks like SVG generation.
Warp’s Productivity dashboard is an eng manager’s dream
Warp's new productivity dashboard centralizes engineering team visibility for managers and CTOs, enabling real-time monitoring of development velocity, code quality improvements, and cost efficiency across distributed teams rather than relying on individual local setups.
She uses Claude Code to find a house with the right vibe
Hillary shares an AI workflow she uses for house hunting that leverages Claude Code to find listings and filter them against her subjective "vibe-specific criteria," which she describes as evaluating whether a space would feel good to wake up in and spend time at.