Best AI model for coding: Fable vs Opus 5 vs Sol vs Grok vs DeepSeek | DHH and Lex Fridman
DHH discusses testing multiple AI models (Fable, Opus 5, Grok, DeepSeek, Claude) on a complex coding task—translating a Python library to Rust—and shares results comparing speed, cost, and capability. He finds Fable to be the strongest model overall, though competitors like Grok 4.6 and DeepSeek offer compelling cost-to-performance tradeoffs.
Summary
DHH explains that the AI coding landscape is highly competitive, with multiple models approaching frontier capability. He conducted a comprehensive test translating the terminal text effects Python library into Rust across multiple AI models to compare real-world performance. Fable completed the task in under 45 minutes, creating a detailed plan and delivering a pixel-perfect Rust implementation that reduced startup time from 86ms to 2ms, achieved 9.6x speedup in execution, and produced a 3MB executable. When Fable ran out of tokens, Opus 5 seamlessly continued and finished the job, demonstrating the value of Fable's upfront planning. DHH then tested other models: Claude 3.5 Sonnet (using Fable's plan) completed it in 1.5 hours at $46 worth of tokens via subscription; GPT Luna failed entirely; Grok 4.6 succeeded in completing the task at $55 per-token cost with 10x speedup; Claude 3 Haiku took a very long time; DeepSeek v4 Flash failed, but DeepSeek Pro succeeded in 2 hours 45 minutes at $23 cost. The results show Fable as fastest but expensive at $550, while Grok and DeepSeek offered similar outputs at 1/10 and 1/20 the cost respectively, with longer execution times. DHH also mentions achieving a 46x execution improvement over the original through multiple agentic research runs. He concludes with a best practice: using Fable for planning and review, then Opus 5 or other models for implementation, and having multiple models cross-check each other's work.
Key Insights
- Fable created a detailed 8-step plan without being asked, which enabled Claude Opus to seamlessly take over mid-task when tokens ran out, demonstrating that upfront planning by stronger models has downstream benefits for task handoffs.
- Grok 4.6 completed a complex Rust translation task successfully at $55 cost with 10x speedup, showing that Xai's model represents a genuine competitive breakthrough despite DHH's earlier skepticism about Grok 4.5.
- DeepSeek Pro achieved the same coding task outcome as Fable in 2 hours 45 minutes at just $23 cost—a 1/20 cost reduction—demonstrating that Chinese open-weight models have reached frontier capability while maintaining dramatic cost advantages.
- A single task that would require 9 months of manual Rust learning by a developer costs $500 to complete via AI agents, making the ROI compelling enough that DHH considered it a steal despite the high absolute cost.
- DHH's standard operating procedure is to have Fable do planning and review work, use Opus 5 for implementation, and employ cross-model verification with other frontier models to catch mistakes, since even the best model (Fable) makes errors.
Topics
Transcript
[0:02] You now run agents all day. What stands out to you as the better as the better model? Who's currently winning? It seems to be changing constantly. So, Fable, Grok-4-6, Gemini, the Chinese open weight models. >> What's amazing to me is that you could rattle off so many different contenders. That this market is so wide open that it does actually change back and forth. That we have real competition. That there are so many labs [0:35] that are able to get either to the frontier or close to it. That's by the way is is remarkable. I still don't fully understand that. But to answer your question, the best model in general right now is Fable. The…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Lex Clips
How Sigmund Freud revolutionized psychiatry | Andrew Scull and Lex Fridman
Andrew Scull discusses how Sigmund Freud and psychoanalysis emerged in late 19th-century America alongside religious healing movements, eventually gaining popularity among intellectuals and artists in the 1920s, despite mainstream psychiatry's resistance to talk therapy as treatment for mental illness.
Do antipsychotic drugs work? | Andrew Scull and Lex Fridman
Andrew Scull discusses the efficacy and serious side effects of antipsychotic drugs, explaining that while they help some patients, many are non-responders, and the medications carry substantial risks including tardive dyskinesia, weight gain, metabolic syndrome, and cognitive dulling. The landmark CATIE study revealed that newer, expensive second-generation antipsychotics are no more effective than older first-generation drugs, with 67-82% of patients dropping out due to inefficacy or intolerable side effects.
Surprising origin of antipsychotic drugs | Andrew Scull and Lex Fridman
Andrew Scull explains how chlorpromazine, an antihistamine synthesized in the 1880s, was accidentally discovered to have psychiatric applications in the 1950s through French Navy lieutenant Alain Laborit's experiments. The drug's rapid adoption was driven by pharmaceutical companies' aggressive marketing to politicians and hospital administrators rather than by psychiatrists' initial enthusiasm, transforming it from a 'major tranquilizer' into the first 'antipsychotic' drug.
Cognitive Behavioral Therapy (CBT) vs Psychoanalysis | Andrew Scull and Lex Fridman
Andrew Scull discusses how clinical psychology and cognitive behavioral therapy (CBT) emerged after WWII as alternatives to psychoanalysis, becoming dominant through better empirical evidence, shorter treatment duration, and federal funding advantages. CBT's symptom-focused approach proved more measurable and reproducible than psychoanalytic theory, though its effectiveness remains limited for serious mental disorders.
Do anti-depressant drugs work? | Andrew Scull and Lex Fridman
Andrew Scull discusses the history and efficacy of antidepressants, particularly SSRIs like Prozac, explaining that while they statistically outperform placebo, the clinical improvement is often marginal and comes with significant side effects including sexual dysfunction and withdrawal difficulties. He also addresses 'diagnostic creep' in psychiatry, where conditions become increasingly broadly defined and diagnosed over time.