TechnicalDiscussion

Best AI model for coding: Fable vs Opus 5 vs Sol vs Grok vs DeepSeek | DHH and Lex Fridman

Lex Clips

DHH discusses testing multiple AI models (Fable, Opus 5, Grok, DeepSeek, Claude) on a complex coding task—translating a Python library to Rust—and shares results comparing speed, cost, and capability. He finds Fable to be the strongest model overall, though competitors like Grok 4.6 and DeepSeek offer compelling cost-to-performance tradeoffs.

Summary

DHH explains that the AI coding landscape is highly competitive, with multiple models approaching frontier capability. He conducted a comprehensive test translating the terminal text effects Python library into Rust across multiple AI models to compare real-world performance. Fable completed the task in under 45 minutes, creating a detailed plan and delivering a pixel-perfect Rust implementation that reduced startup time from 86ms to 2ms, achieved 9.6x speedup in execution, and produced a 3MB executable. When Fable ran out of tokens, Opus 5 seamlessly continued and finished the job, demonstrating the value of Fable's upfront planning. DHH then tested other models: Claude 3.5 Sonnet (using Fable's plan) completed it in 1.5 hours at $46 worth of tokens via subscription; GPT Luna failed entirely; Grok 4.6 succeeded in completing the task at $55 per-token cost with 10x speedup; Claude 3 Haiku took a very long time; DeepSeek v4 Flash failed, but DeepSeek Pro succeeded in 2 hours 45 minutes at $23 cost. The results show Fable as fastest but expensive at $550, while Grok and DeepSeek offered similar outputs at 1/10 and 1/20 the cost respectively, with longer execution times. DHH also mentions achieving a 46x execution improvement over the original through multiple agentic research runs. He concludes with a best practice: using Fable for planning and review, then Opus 5 or other models for implementation, and having multiple models cross-check each other's work.

Key Insights

  • Fable created a detailed 8-step plan without being asked, which enabled Claude Opus to seamlessly take over mid-task when tokens ran out, demonstrating that upfront planning by stronger models has downstream benefits for task handoffs.
  • Grok 4.6 completed a complex Rust translation task successfully at $55 cost with 10x speedup, showing that Xai's model represents a genuine competitive breakthrough despite DHH's earlier skepticism about Grok 4.5.
  • DeepSeek Pro achieved the same coding task outcome as Fable in 2 hours 45 minutes at just $23 cost—a 1/20 cost reduction—demonstrating that Chinese open-weight models have reached frontier capability while maintaining dramatic cost advantages.
  • A single task that would require 9 months of manual Rust learning by a developer costs $500 to complete via AI agents, making the ROI compelling enough that DHH considered it a steal despite the high absolute cost.
  • DHH's standard operating procedure is to have Fable do planning and review work, use Opus 5 for implementation, and employ cross-model verification with other frontier models to catch mistakes, since even the best model (Fable) makes errors.

Topics

AI model comparison and benchmarkingAgentic coding and code generation capabilitiesCost-to-performance tradeoffs in AI modelsPlanning and task decomposition in AI agentsReal-world productivity gains from AI coding tools

Transcript

[0:02] You now run agents all day. What stands out to you as the better as the better model? Who's currently winning? It seems to be changing constantly. So, Fable, Grok-4-6, Gemini, the Chinese open weight models. >> What's amazing to me is that you could rattle off so many different contenders. That this market is so wide open that it does actually change back and forth. That we have real competition. That there are so many labs [0:35] that are able to get either to the frontier or close to it. That's by the way is is remarkable. I still don't fully understand that. But to answer your question, the best model in general right now is Fable. The…

Full transcript available for MurmurCast members

Sign Up to Access

More from Lex Clips

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.