What If We Stopped Using GPUs? | YC Paper Club
A YC Paper Club discussion exploring alternatives to GPU-based computing, featuring presentations on optical computing, neuromorphic chips, and biological computing approaches. Speakers argue that GPU efficiency (FLOPS per joule) has plateaued, necessitating fundamentally different computational substrates inspired by the brain's 20-watt energy efficiency.
Summary
The event begins with François, the moderator, introducing the core problem: GPU efficiency improvements have stalled over the past two years despite exponential growth in computational demands from transformer models. He explains that since 2020, NVIDIA shifted focus from energy efficiency to memory bandwidth and capacity due to attention's quadratic memory requirements. François proposes exploring alternative computing paradigms, particularly those inspired by the brain's remarkable 20-watt power consumption versus modern AI systems' massive energy demands.
François presents his own research on SPSA (Simultaneous Perturbation Stochastic Approximation), a gradient-free optimization method that uses finite differences rather than backpropagation. He argues this approach scales better with model parallelization and proposes 'Soma' (Shadow Mixture Optimization Collections), clustering datasets and distributing small expert models across GPUs worldwide. The key insight is that gradient estimation noise becomes worse with model size, but decreases when models are split—suggesting massive parallelization could work better without backpropagation.
Ilker presents optical computing, arguing that photons have 10,000 times lower losses than electrons over distance and massively parallel properties since photons don't interact. However, he highlights critical challenges: converting digital data to analog/optical and back wastes energy, passive optics struggle with nonlinear activation functions, and programming/calibration adds overhead. His team's work on diffusion models shows optical systems can achieve comparable results to digital networks with better energy efficiency for specific tasks, and they demonstrate power-law scaling benefits when adding optical parameters.
Alok discusses neuromorphic computing—using brain-inspired principles in chip design. He emphasizes the brain's remarkable co-adaptation of hardware and software over evolution, and notes that modern neural networks have strayed far from biological principles by adopting backpropagation and transformers. He categorizes neuromorphic research into three areas: energy efficiency through analog learning, event-driven computation using spike-based networks, and exotic substrates (photonics, memristors, coupled oscillators). Alok notes the field remains in research and development, with DMatrix showing promise by combining SRAM with logic in digital format.
Sean presents work on training biological neural cultures from rat brains to play Doom, treating the system as a black-box input-output dynamic system. Rather than manually encoding/decoding, they use PPO (Proximal Policy Optimization) with learned encoders that translate game states into electrode stimulation frequencies/amplitudes, and decoders that read spikes as actions. Feedback is provided through asynchronous (negative) and synchronous (positive) stimulation patterns, modulated by TD error. The work demonstrates that biological systems can learn complex tasks when given proper learning algorithms, though scalability to human-level intelligence remains unresolved.
Key Insights
- François argues that FLOPS per joule improvements have stalled over the past two years despite increasing model size, indicating GPU-based approaches are hitting physical efficiency limits that require entirely new computational substrates.
- SPSA (Simultaneous Perturbation Stochastic Approximation) gradient-free optimization becomes more scalable with model parallelization because gradient estimation noise is limited per expert size, requiring only ~64 disturbances per step rather than full backpropagation.
- Ilker demonstrates that optical computing achieves energy advantages for specific tasks like diffusion-based image generation, but the primary bottleneck is not computation itself—it's the overhead of converting between digital, analog, and optical domains for data storage and retrieval.
- Alok argues that the brain represents the best model of hardware-software co-design ever evolved, but modern deep learning has departed from biological principles by adopting backpropagation and transformers, suggesting future neuromorphic gains require careful selection of which brain principles actually apply to silicon.
- Sean's work training biological neural cultures shows that with proper learning algorithms (PPO with learned encoders/decoders), biological systems can learn complex tasks, but the critical insight is that oversized decoders can fool performance metrics—the biological substrate itself must genuinely learn.
Topics
Transcript
[0:07] Okay, please. Welcome to the Club alternative calculations. So, this one The evening began with dinner with mine a close friend who gave me this topic . I thought that was it. really interesting topic, which you guys can to take with you on your dinner. It was like this: if you could put aliens only one question, this super smart aliens, one question you would like to ask did you put it? So we we think about you know how you are solve the problem about the power law system? And I thought that we seem to know the areas Dyson or something similar about nuclear [0:38] systems. But I I thought it was more interesting. would be…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Y Combinator
A camera that can see through walls
Bill Abster from Applied Electromagnetics discusses his company's breakthrough technology that uses radio waves and high-frequency antennas to create 3D images of objects behind opaque materials. The system processes echoes in real-time using GPUs to generate detailed depth maps, and the company has achieved significant progress in just 4 months by maintaining rapid iteration cycles.
Robotic factories that build robots
Tensor is building a 12,000 sq ft robotic factory that manufactures robots, including humanoids and specialized grippers. The company, founded by Berkeley researchers with autonomous vehicle and manufacturing backgrounds, treats the entire factory as a learning system that improves manufacturing processes with each customer order.
New LLMs Are Unlocking Robot-Use Agents
This discussion explores how large language models like Claude (Astra) are enabling robot control through code generation and tool use, building on foundational work like RT2 and visual language-action models. The speakers from Waddle Labs and RoboCurve explain how models trained on diverse data modalities—including web text, computer usage, and coding—are converging toward general-purpose robot agents that can perform complex physical tasks within the next 2 years.
Rocket cargo delivery anywhere on Earth in minutes
Matthew, CEO of Hover Arrow, discusses the company's mission to deliver cargo anywhere on Earth within minutes using rocket technology. The startup has secured $1.37 billion in commercial letters of intent, with primary focus on military applications and emerging interest from industries requiring rapid, temperature-controlled delivery of perishable goods.
Autonomous construction on Earth and beyond
James, co-founder of Fossma Robotics, discusses building autonomous construction equipment for large-scale infrastructure on Earth with the long-term goal of enabling construction on Mars and the moon. The company has already deployed systems that installed 11,000 solar panels and recently signed a major contract with the top construction company in the US.