Why smarter AI models could drive up compute prices 10x
As AI models become more capable and valuable, compute capacity growth (3x yearly) cannot keep pace with revenue growth (10x yearly), forcing labs to either increase margins, raise compute prices, or shift more resources to inference. The speaker argues compute prices will likely increase significantly as AI models approach human-level capabilities, making compute a scarce resource similar to skilled labor.
Summary
The speaker analyzes the economic tension facing AI labs like Anthropic that are experiencing 10x year-over-year revenue growth while compute capacity only increases 3x annually. Using Anthropic as a case study (growing from $9B to $100-150B projected revenue), the speaker identifies three mechanisms labs are using to bridge this gap: increasing margins (inference margins reportedly jumped from 40% to 80%), raising compute prices (spot prices up 40% since February), and shifting compute allocation toward inference rather than training.
The speaker argues that labs will likely continue prioritizing training over inference despite margin pressures, as they view inference revenue primarily as a vehicle to attract investment for the next generation of models. This leaves two primary paths: either margins compress to unsustainable levels (unlikely above 90% in competitive markets) or compute prices must rise substantially.
The core argument centers on compute scarcity and value realization. As AI models become human-level capable, they can monetize compute at rates comparable to highly-paid professionals—suggesting H100 GPUs could theoretically rent for over $250K annually (15x current spot prices) if they replicated a software engineer's productivity. The speaker cites Google's arrangement with SpaceX, where Google pays 2x spot prices for guaranteed capacity, as evidence this dynamic is already emerging.
The speaker acknowledges the uncertainty about whether labor supply shocks would suppress value (following standard economics) or whether the shock will be too fast for market adjustment. They also note three structural limits to continued 3x annual compute scaling: Moore's Law contribution (1.4x) is plateauing, fab capacity expansion (1.2x) is constrained by ASML EUV machine production, and AI's absorption of wafer allocation from consumer devices (1.8x) will saturate soon as AI approaches 86% of leading-edge TSMC capacity.
Finally, the speaker expresses concern about drawing false parallels to past scarcity predictions (Simon-Ehrlich bet) while arguing that compute supply is fundamentally less elastic than commodity markets, making sustained price increases more likely in the pre-singularity period.
About this episode
<p>This is a video recording of a post I wrote last week. If you want to read the original you can check it out <a href="https://www.dwarkesh.com/p/why-compute-might-get-10x-more-expensive" target="_blank">here</a>.</p><p>Thanks to <a href="https://mercury.com/command" target="_blank">Mercury</a> for sponsoring this video. <a href="https://mercury.com/command" target="_blank">Mercury’s</a> built-in AI, Command, helps me close my books and saves me a bunch of time. At the end of each month, Command categorizes my transactions and provides its rationale for every choice: I just review, fix anything that’s off, and approve... and then Mercury syncs everything to QuickBooks. Get started at <a href="https://mercury.com/command" target="_blank">mercury.com/command</a></p> <br /><br />Get full access to Dwarkesh Podcast at <a href="https://www.dwarkesh.com/subscribe?utm_medium=podcast&utm_campaign=CTA_4">www.dwarkesh.com/subscribe</a>
Key Insights
- AI labs are experiencing a structural gap where revenue grows 10x yearly while compute capacity only grows 3x, forcing them to choose between raising prices, increasing margins above sustainable levels, or shifting resources toward inference at the cost of future capability development.
- The speaker argues that as AI models approach human-level capability, compute will be valued at rates comparable to skilled professionals—potentially 15x current prices—because the same compute could replace $250K+ annual software engineer salaries.
- Three components of annual compute scaling (Moore's Law at 1.4x, fab expansion at 1.2x, and wafer reallocation from consumer devices at 1.8x) are each approaching natural limits, making continued 3x growth difficult to sustain beyond the next few years.
- Labs deliberately avoid prioritizing inference work despite margin benefits because they view inference revenue as primarily a funding mechanism for training the next generation of models, not as their core business.
- Google's agreement to pay 2x spot prices for guaranteed compute capacity from SpaceX demonstrates that frontier labs are already willing to accept significant price premiums for secure, large-scale access—evidence that the compute price increase scenario is actively materializing.
Topics
Transcript
Today I want to talk about what the compute situation for the labs will look like over the next few years. For the last three consecutive years, Anthropix revenue has 10x year over year, and it's likely to do so again this year. So they ended last year with nine billion in revenue. I think they'll probably end this year with somewhere between 100 billion to $150 billion in revenue. Now, for this trend to continue, Anthropix would need to make $1 trillion in revenue by the end of next year. Of course, there's no deep reason why this has to be true. It's a very wild conclusion and it's ultimately a question of next year. Of course, there's no…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Dwarkesh Podcast
Adam Brown – Einstein's happiest thought: General Relativity from scratch
Adam Brown explains Einstein's General Relativity from first principles, showing how the equivalence of inertial and gravitational mass led Einstein to conceptualize gravity as curved spacetime rather than a force. The lecture progresses from special relativity through the geometric interpretation of gravity to black holes, demonstrating GR's explanatory power across vastly different scales.
Grant Sanderson – AI and the future of math
The discussion centers on the rapid advancements of AI in mathematics, exploring its implications for the future of math and related fields. The conversation highlights how AI's capabilities impact traditional mathematical roles, the process of knowledge creation, and the potential for new insights in various domains.
The next big breakthrough will be AIs learning on the job
The speaker discusses how AI labs are betting on reinforcement learning from verified rollouts (RLVR) to achieve AGI, but argues this approach has fundamental limitations. He contends that true general intelligence requires continual on-the-job learning through weight updates, which current scaling paradigms don't adequately address.
The data black hole at the center of AI
The transcript argues that AI's primary driver of progress is data quantity and quality rather than architectural improvements or scaling, highlighting a massive gap in sample efficiency between humans and AI models. The speaker contends that current AI systems are fundamentally different from human intelligence, requiring orders of magnitude more data to learn skills. Despite this inefficiency, AI can still automate white-collar work due to the economics of scale and parallelism.
Ada Palmer – Machiavelli is the most misunderstood thinker of all time
Ada Palmer discusses Machiavelli's political theories and their historical context, emphasizing the instability of Italian city-states and the influence of the papacy. She explores how Machiavelli's personal experiences and insights shaped his writings, particularly in 'The Prince' and 'Discourses on Livy'.