TechnicalDiscussion

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Neil Movva, founder of SAIL Research, discusses building a 'token factory' designed for long-running AI agents by optimizing across the full stack—software, hardware, and power—to achieve unbeatable cost per token. He argues the future of AI inference will shift from low-latency chatbots to background agents running for hours or days, fundamentally changing GPU optimization priorities from speed to throughput.

Summary

Neil Movva founded SAIL Research to build the cheapest possible inference infrastructure for AI tokens, specifically targeting long-running agent workloads rather than real-time chatbot interactions. The company's strategy is three-pronged: software optimization through GPU kernels and batching efficiency, hardware arbitrage across multiple chip types and providers, and creative power sourcing including renewable energy and distributed small data centers.

Movva argues that the AI inference market is experiencing a fundamental shift. Whereas the past year was dominated by companies like Cursor optimizing for low-latency responses, the future belongs to proactive agents that operate in the background for hours, days, or weeks without human-in-the-loop interaction. This changes everything about hardware optimization: GPUs have an inherent tradeoff between latency (narrow, fast) and throughput (wide, slow). NVIDIA optimizes for low latency through technologies like NVLink, which reduces minimum latency but requires 8x more hardware for sublinear speedup. For background agents, throughput optimization is more valuable than latency optimization, creating an opening for alternative chip strategies.

Movva spent his early career at NVIDIA learning about Tensor Cores, GPU efficiency, and the company's obsessive culture of chasing 'speed of light'—pushing hardware to its theoretical maximum performance. He applies this performance engineering mindset to SAIL, but redirected toward throughput rather than latency. He argues current systems waste enormous amounts of compute: NVIDIA ships 5 million Blackwell chips annually, many sitting idle in private pools or warehouses because compute isn't orchestrated efficiently as a shared global resource.

On the hardware side, Movva advocates for a 'scavenger strategy': buying any chip, from any vendor, at any price point, wherever it exists globally. AMD chips are undervalued because perceptions lag reality; newer players like Cerebrus and Grok make interesting bets on different memory hierarchies (SRAM vs HBM) but work best as accelerators paired with traditional GPUs. NVIDIA's dominance will persist, but not because competitors can't catch up—rather because NVIDIA maintains balanced choices that work adequately across scenarios, while competitors can be spiky. For inference specifically, the HBM shortage constrains NVIDIA more than compute, creating opportunity elsewhere.

On data centers, Movva rejects the conventional wisdom that you need large, redundant, 10+ megawatt facilities. Instead, he's building 1-megawatt distributed data centers across the US—about 8 racks each—with minimal redundancy, accepting 95% uptime because background agents tolerate occasional latency spikes. This distributes risk, avoids power density bottlenecks, and opens access to cheaper power sources.

On energy, Movva's most contrarian position is willingness to use intermittent renewable power (solar and wind). Whereas enterprises need consistent baseload power, he can tolerate data center outages measured in days or weeks by predictively shifting workloads globally based on weather patterns. This gives him access to power no one else wants. He's passionate about proving that you can achieve radical efficiency by reconfiguring the entire system—from chip selection, to software stacks, to power sources—in concert.

On open vs. closed models, Movva argues the premium frontier labs (Anthropic, OpenAI) charge for being 3-6 months ahead will shrink as distillation becomes inevitable and latent. Code from closed-source models already trains the next generation of open models through GitHub. He's bullish on open source long-term and believes abundant, cheap inference is the key unlock—not proprietary models.

On future improvements, Movva identifies KV cache compression as the biggest unsolved efficiency problem. Modern systems store kilobytes per token; an order of magnitude improvement is possible. He also emphasizes the fundamental importance of hiring people driven by curiosity about performance engineering rather than specific tool expertise, and building a culture of collaborative learning.

About this episode

My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time.  In that world, latency matters less and cost matters more, and Neil has built the whole company around driving the cost of a token as low as it can possibly go. What makes this conversation special is that it is one of the most detailed tours I have ever done through the full stack of intelligence, the software, the chips, and the power, and how all three connect.  Along the way we cover the trade-off between speed and cost that lives inside every GPU, his scavenger strategy for buying the chips and power nobody else wants, his contrarian view on Nvidia, and why the premium the frontier labs charge for being three to six months ahead may not last.  Please enjoy my conversation with Neil Movva. For the full show notes, transcript, and links to mentioned content, check out the episode page ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠here⁠⁠⁠⁠⁠.  ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at ⁠colossus.com/subscribe⁠. ----- ⁠Ramp’s⁠ mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ ⁠ramp.com/invest⁠⁠ to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, ⁠Vanta⁠ continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to ⁠vanta.com/invest⁠.  ----- WorkOS⁠ is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- ⁠Ridgeline⁠ has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ridgeline.ai⁠. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:20) Neil Movva (00:03:22) Building a Token Factory (00:05:32) The Rise of Long-Running Agents (00:08:47) Deep Research and Cybersecurity (00:15:12) The Full Stack of Intelligence (00:20:03) Throughput Versus Latency (00:24:58) The Future of AI Chips (00:33:19) Why Transformers Work (00:36:43) The Future of Data (00:44:05) The Market for AI Chips (00:47:56) Is the AI Boom Different? (00:51:08) Reinventing the Data Center (00:56:43) Scavenging Power (01:01:04) Where Compute Is Most Inefficient (01:07:02) Open Versus Closed Models (01:10:37) A Trillion Tokens a Day (01:12:42) The Contrarian Case on NVIDIA (01:14:38) Advice for AI Hardware Founders

Key Insights

  • Movva claims the defining shift in AI infrastructure is moving from low-latency chatbot responses to long-running background agents, which fundamentally changes optimal hardware choices from latency-minimized (NVIDIA's current focus) to throughput-maximized systems.
  • He argues that GPUs have an unbreakable hardware tradeoff between latency and throughput: serving one user quickly requires different parallelism schemes than batching many users' work, analogous to choosing between a private car and a bus.
  • Movva contends that NVIDIA's obsession with NVLink and low-latency optimization is excellent for real-time inference but wasteful for background workloads, creating an opening for alternative chips without advanced interconnects if coupled with different parallelism strategies.
  • He claims that the world is extremely inefficient at compute utilization, with millions of GPUs sitting idle in private company pools or warehouses, representing wasted silicon and power that better orchestration could recapture.
  • Movva argues that AMD chips are undervalued due to perception lag and that custom kernel optimization can extract significantly more performance from non-NVIDIA hardware, though NVIDIA maintains an advantage through consistency across use cases.
  • He contends that KV cache compression represents the largest unsolved efficiency frontier in inference, with current systems storing an order of magnitude more data per token than theoretically necessary.
  • Movva claims distributed 1-megawatt data centers with ~95% uptime are economically superior to traditional 10+ megawatt facilities for inference, because background agents tolerate occasional latency spikes from outages.
  • He argues that renewable and intermittent power sources (solar, wind) are viable for compute at scale if workloads are flexible and globally distributed, enabling access to cheaper power competitors cannot profitably use.
  • Movva contends that the premium frontier labs command for being 3-6 months ahead of open source models is sustainable but not permanent, as distillation and latent knowledge transfer through user artifacts makes capability diffusion inevitable.
  • He claims that the current consensus around TSMC's superiority over Western fabs is exaggerated; while process gaps exist, they are 2x performance-per-watt at worst, not the 10x difference geopolitical discourse suggests.
  • Movva argues that the most important hire for performance engineering is someone driven by intrinsic curiosity about hardware behavior, not domain expertise in CUDA or specific tools, because frameworks evolve faster than fundamental intuition.
  • He contends that abundant, cheap inference unlocks entirely new product categories and use cases that expensive inference cannot serve, justifying massive investment in multi-layer stack optimization and global resource arbitrage.

Topics

Long-running AI agents vs. real-time inference latencyGPU throughput vs. latency tradeoffsHardware arbitrage and multi-chip optimizationNVIDIA dominance and competitive chip alternativesDistributed small data centers vs. large monolithic facilitiesRenewable and intermittent power for computeKV cache compression and efficiencyOpen source vs. closed-source model dynamicsKernel and software optimization for peak GPU efficiencyGlobal compute orchestration and resource allocationTSMC and advanced packaging bottlenecksTest-time compute scaling for agent reasoning

Transcript

Ramp is the only platform built to make your finance team leaner, faster, and better, saving businesses 5% annually on average so you can stay focused on growth. Ramp customers grow revenue 3.2 times faster than the average American business. Visa, Vercel, Cursor, Stripe, Notion, 11Lab, Shopify, and 70,000 other businesses all now run on Ramp. Mine does too, and so should yours. Learn more at ramp.com slash invest. Mine does too, and so should yours. Learn more at ramp.com slash invest. Felix by Rogo is a personal finance agent that turns a single prompt into finished client-ready work using your firm's own templates, context, and standards. Send Felix an email like, "'Take these comments and turn them for me,'…

Full transcript available for MurmurCast members

Sign Up to Access

More from Invest Like the Best with Patrick O'Shaughnessy

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.