DiscussionTechnical

Robot-Use Agents: Why General-Purpose Models May Win in Robotics

Y Combinator Startup Podcast29m 49s

General-purpose language models like Astra are emerging as effective robot control agents by leveraging their learned representations of the physical world, moving away from task-specific robotics models toward a unified approach where pre-training on diverse data (code, computer use, vision) enables strong spatial reasoning applicable to robot manipulation tasks.

Summary

The discussion explores how large language models are becoming viable for robot control, challenging the traditional approach of building specialized robotics foundation models. Speakers from Waddle Labs and RoboCurve discuss the evolution from RT2 (Vision Language Action models that required fine-tuning) to current frontier models like Astra that can control robots zero-shot by writing code or issuing tool calls.

The key shift is from direct action prediction (outputting coordinates) to code generation as policies, where models write Python code that orchestrates robot control. This approach offers significant advantages: models can allocate more compute for complex tasks, leverage tool use and reflection, and benefit from in-context learning. Early work like the Voyager paper demonstrated that coding agents could create and invoke tools on-the-fly, compressing experience into reusable functions.

The conversation addresses different learning paradigms: in-context learning (ICL) is cheap but limited by context window and model training; low-rank adapters (LoRA) offer intermediate efficiency; and full fine-tuning scales to large datasets. The speakers argue for a hierarchical approach where general foundation models generate skills and memories that are consolidated into faster, specialized policies through distillation.

A central theoretical framework is the Platonic Representation Hypothesis, suggesting that sufficiently capable models trained on different modalities converge to similar internal representations of the world. This implies that strong language models automatically learn strong representations for robotics without explicit robot-specific training. Astra's improvements stem from extensive pre-training on computer use data (CAD, Blender interaction), egocentric video, and diverse vision data, which teaches spatial reasoning transferable to physical robot manipulation.

Practical demonstrations show these models performing tasks like block manipulation, uncapping bottles, and multi-robot coordination. The Waddle Labs harness concept encapsulates learned skills into reusable code libraries, while managing latency through skill consolidation. The speakers predict general-purpose robots capable of executing natural language instructions at human-competent levels within two years, presenting both technological and societal implications.

About this episode

<p>One of the biggest surprises in AI over the last few years has been how well coding agents generalize beyond software. In a recent essay, MIT professor Philip Isola argued that we may be entering the era of robot-use agents: general-purpose models that can control different robots, write policies, and learn new physical tasks with little or no robot-specific training.In this episode of Decoded, we're joined by the founders of Waddle Labs and RoboCurve, two of the startups whose work helped drive this realization. They're working at the frontier of using general-purpose models to control robots, and their recent demos helped inspire the growing conversation around robot-use agents. </p><p>Together, we dig into the research behind that idea, from code-as-policies and vision-language-action models to the harnesses and evals needed to make these systems work in the real world.</p>

Key Insights

  • Frontier language models achieve robot control without task-specific fine-tuning by leveraging pre-training on diverse data including computer use and CAD data, which teaches spatial reasoning transferable to physical manipulation.
  • Code generation enables more efficient scaling of computational resources than direct action prediction because models can allocate additional compute through tool use and reflection when encountering complex tasks.
  • In-context learning saturates after 20-40 examples and is bounded by the model's context window from pre-training, necessitating skill consolidation and distillation strategies for sustained improvement and practical deployment.
  • Computer interface design (GUIs, file systems, 3D manipulation tools) inadvertently created training data formats that teach spatial reasoning comparable to the physical world, making computer use data surprisingly effective for robot learning.
  • Speakers predict consensus among frontier labs that general-purpose robots executing natural language instructions at human-competent levels will emerge within two years, representing a comparable inflection point to ChatGPT but for physical capabilities.

Topics

General-purpose language models for robot controlCode generation as policies versus direct action predictionIn-context learning and knowledge consolidationPlatonic Representation Hypothesis and cross-modal transferRobot harness systems and skill librariesPre-training data diversity and spatial reasoningVision Language Action models evolution

Transcript

One of the big surprises the last few years has been the generalizability of coding agents across different domains. And now, frontier researchers are showing that this includes controlling robots. This led MIT professor Philip Isola to suggest in a recent viral essay that we may be entering the era of robot use agents, where general purpose models could make different robots more capable. So today, Francois and I invited the founders of Waddle Labs and RoboCurve, two groups of startups that are working at the very frontier of making robots more capable with LLMs. Maybe you guys want to briefly introduce yourselves and just say a little bit about what each of your companies focuses on. I'm Hanmei. I'm…

Full transcript available for MurmurCast members

Sign Up to Access

More from Y Combinator Startup Podcast

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.