The MAD Podcast with Matt Turck
MurmurCast publishes AI-generated summaries of The MAD Podcast with Matt Turck’s YouTube episodes — 24 summarized so far, covering Open-source AI ecosystem and community collaboration, NVIDIA NeMoTron family of models (Nano, Super, Ultra), Hybrid SSM-Transformer architecture, Mixture of Experts (MoE) and sparse computation, 4-bit arithmetic pre-training (NVF P4), Multi-token prediction and inference optimization. Each summary distills the key insights, topics, and takeaways so you can decide what’s worth your time before pressing play.
NVIDIA’s Bryan Catanzaro: Why More Compute Isn’t Enough
Bryan Catanzaro, who leads NVIDIA's NeMoTron frontier AI models, discusses how open-source AI is accelerating through community collaboration, explains the technical innovations in NeMoTron 3 (hybrid SSM-transformer architecture, mixture of experts, multi-token prediction), and argues that open technologies are safer and more aligned with how effective organizations actually work.
Cloudflare CEO: Bot Takeover, Edge AI & The Hard Decision Every CEO Will Face
Matthew Prince, CEO of Cloudflare, discusses how bot traffic has surpassed human traffic on the internet as of mid-2026, driven by AI agents and LLMs. He explores how this fundamental shift is forcing a reimagining of internet infrastructure, business models, and organizational structures, with Cloudflare positioned at the center of these changes through products like Workers, AI Gateway, and edge computing solutions.
Why Idle GPUs Bleed Cloud Companies Dry #ai #podcast
The podcast discusses how GPU depreciation costs are the largest component of cloud computing expenses, and that GPU utilization directly impacts per-hour costs. Cloud companies gain competitive advantage by building beloved products that drive high GPU utilization rates.
The Physics of an AI Token #ai #podcast
The transcript explains the energy-to-computation pipeline for AI systems, tracing how raw energy sources (photons or natural gas) are converted through power plants into electrical power, then processed by servers into floating-point operations, and finally transformed into AI tokens per second.