Robot-Use Agents: Why General-Purpose Models May Win in Robotics
General-purpose language models like Astra are emerging as effective robot control agents by leveraging their learned representations of the physical world, moving away from task-specific robotics models toward a unified approach where pre-training on diverse data (code, computer use, vision) enables strong spatial reasoning applicable to robot manipulation tasks.
Summary
The discussion explores how large language models are becoming viable for robot control, challenging the traditional approach of building specialized robotics foundation models. Speakers from Waddle Labs and RoboCurve discuss the evolution from RT2 (Vision Language Action models that required fine-tuning) to current frontier models like Astra that can control robots zero-shot by writing code or issuing tool calls.
The key shift is from direct action prediction (outputting coordinates) to code generation as policies, where models write Python code that orchestrates robot control. This approach offers significant advantages: models can allocate more compute for complex tasks, leverage tool use and reflection, and benefit from in-context learning. Early work like the Voyager paper demonstrated that coding agents could create and invoke tools on-the-fly, compressing experience into reusable functions.
The conversation addresses different learning paradigms: in-context learning (ICL) is cheap but limited by context window and model training; low-rank adapters (LoRA) offer intermediate efficiency; and full fine-tuning scales to large datasets. The speakers argue for a hierarchical approach where general foundation models generate skills and memories that are consolidated into faster, specialized policies through distillation.
A central theoretical framework is the Platonic Representation Hypothesis, suggesting that sufficiently capable models trained on different modalities converge to similar internal representations of the world. This implies that strong language models automatically learn strong representations for robotics without explicit robot-specific training. Astra's improvements stem from extensive pre-training on computer use data (CAD, Blender interaction), egocentric video, and diverse vision data, which teaches spatial reasoning transferable to physical robot manipulation.
Practical demonstrations show these models performing tasks like block manipulation, uncapping bottles, and multi-robot coordination. The Waddle Labs harness concept encapsulates learned skills into reusable code libraries, while managing latency through skill consolidation. The speakers predict general-purpose robots capable of executing natural language instructions at human-competent levels within two years, presenting both technological and societal implications.
About this episode
<p>One of the biggest surprises in AI over the last few years has been how well coding agents generalize beyond software. In a recent essay, MIT professor Philip Isola argued that we may be entering the era of robot-use agents: general-purpose models that can control different robots, write policies, and learn new physical tasks with little or no robot-specific training.In this episode of Decoded, we're joined by the founders of Waddle Labs and RoboCurve, two of the startups whose work helped drive this realization. They're working at the frontier of using general-purpose models to control robots, and their recent demos helped inspire the growing conversation around robot-use agents. </p><p>Together, we dig into the research behind that idea, from code-as-policies and vision-language-action models to the harnesses and evals needed to make these systems work in the real world.</p>
Key Insights
- Frontier language models achieve robot control without task-specific fine-tuning by leveraging pre-training on diverse data including computer use and CAD data, which teaches spatial reasoning transferable to physical manipulation.
- Code generation enables more efficient scaling of computational resources than direct action prediction because models can allocate additional compute through tool use and reflection when encountering complex tasks.
- In-context learning saturates after 20-40 examples and is bounded by the model's context window from pre-training, necessitating skill consolidation and distillation strategies for sustained improvement and practical deployment.
- Computer interface design (GUIs, file systems, 3D manipulation tools) inadvertently created training data formats that teach spatial reasoning comparable to the physical world, making computer use data surprisingly effective for robot learning.
- Speakers predict consensus among frontier labs that general-purpose robots executing natural language instructions at human-competent levels will emerge within two years, representing a comparable inflection point to ChatGPT but for physical capabilities.
Topics
Transcript
One of the big surprises the last few years has been the generalizability of coding agents across different domains. And now, frontier researchers are showing that this includes controlling robots. This led MIT professor Philip Isola to suggest in a recent viral essay that we may be entering the era of robot use agents, where general purpose models could make different robots more capable. So today, Francois and I invited the founders of Waddle Labs and RoboCurve, two groups of startups that are working at the very frontier of making robots more capable with LLMs. Maybe you guys want to briefly introduce yourselves and just say a little bit about what each of your companies focuses on. I'm Hanmei. I'm…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Y Combinator Startup Podcast
Paul Graham On Startups, Ambition, and Great Founders
Paul Graham discusses YC's evolution over 21 years, arguing that startups funded today are more ambitious than in the past, and explores how founder ambition, shipping speed, and formidability drive success. He shares perspectives on AI's surprising development as a 'bullshitting' system rather than a perfect-to-complex progression, and explains how YC's batch model has remained fundamentally unchanged despite scaling.
The World’s Largest Electric Aircraft Just Flew
HART Aerospace has successfully flown the world's largest electric aircraft, a 100-foot wingspan hybrid-electric plane designed to reduce regional air travel costs. The company, founded by Anders Forslund, evolved from a 3D-printed model seven years ago to a 40-person team building a commercially viable airliner that swaps jet engines for electric motors.
Max Junestrand: You Need The Willingness To Learn Faster Than Anyone Else
Max Junestrand of Legora describes the company's rapid growth from zero to $100M ARR in 18 months by building an AI operating system for lawyers. He discusses the importance of learning the market deeply, maintaining competitive culture, hiring for growth potential rather than prestigious credentials, and betting on improving AI models without fine-tuning.
Susan Kare: Designing Icons & Graphics For the Original Mac
Susan Kare, the iconographer for the original Macintosh, discusses her journey designing system fonts, icons, and graphics under severe technical constraints (16x16 black and white pixels). She shares design principles learned from mentors like Paul Rand, her experiences across multiple tech companies, and the philosophy that simplicity, metaphor, and meaningful design create universal, memorable user interfaces.
Chelsea Finn: This is the State of the Art in Robotics
Chelsea Finn from Physical Intelligence discusses advancing robotics toward general-purpose models that achieve high reliability through reinforcement learning, memory systems, and diverse training data. She demonstrates how robots can perform complex real-world tasks autonomously and argues that physical AI has reached a ChatGPT-like era comparable to language models, with models now being deployed in real-world applications.