The unexpected generalization of coding agents is one of the biggest surprises of recent years. Models trained on millions of lines of internet code gained a problem-solving discipline that spills into domains they never saw. That leap now reaches robot arms, and general-purpose models begin to steer different bodies with one language. For the first time, the rule that every body needs its own brain is seriously shaken.
The Isola thesis and the cloud puppeteer
MIT researcher Phillip Isola names this break clearly in his September 7, 2026 essay. A language model, he argues, can use a robot as a tool the same way it uses a calculator or a computer. According to MIT, every internet-connected robot becomes a potential hand for strong models such as Fable and Astra. In this vision intelligence stays in the cloud, the body's software never changes, and updates arrive over the air.
Waddle Labs founders Jaime and Vincent try to prove this vision on hardware. The team builds a runtime layer so large models can steer robots effectively, collects data, and trains better models from it. According to waddlelabs.ai, the agent splits a goal into subtasks, looks at camera feeds, writes control code, and calls ready action models as tools when needed. The claim is to start work without collecting fresh data for each new arm or gripper.
The RoboCurve side keeps the stage honest through measurement. Co-founder Jay says they test arms, grippers, humanoids, and quadrupeds with equal rigor. According to robocurve.org, the Y Combinator-backed public-benefit company reports frontier capabilities openly. Its open-source Inspect Robots framework aims to replace cherry-picked demo clips with a continuous, standard yardstick.
Short clips that spread in recent weeks fueled the debate. One recording shows a large model loosening a lid, another shows two robots splitting work, a third shows an arm opening a pen. According to The Decoder, the new StationeryBench test gave the spatial leap a hard number: with dual-arm YAM robots over 200 trials, Astra fully finished 7 of 100 tasks while the rival model finished zero. The median progress score of 46 versus 12 turned spatial reasoning into something measurable.
RT2, direct action, and the thinking gap
Google DeepMind introduced RT2 in 2023 as the first serious stop on this road, and its constraint explains what changed. According to DeepMind, the PaLI-X vision-language backbone was fine-tuned on demonstrations collected by 13 robots over 17 months at a kitchen counter. Instead of English, the model emitted end-effector positions that translate into joint commands. Web-scale visual and linguistic pretraining sharply improved generalization to objects and commands missing from the robot data. Yet the model had to produce the action directly in a single pass, like being forced to print the final mark in the GSM8K analogy without showing the working, while today's agents open a chain-of-thought interval, think across turns, and only then emit one command, a flexibility that decides complex tasks.
The intuition matches the bitter-lesson principle: give a model more autonomy and compute and many finely tuned skills emerge on their own. Freeing a strong foundation model beats locking the architecture into robot-only data, sooner and cheaper. The Google Research line behind Code as Policies focuses on the kind of data for exactly this reason. According to research writing, few-shot instruction-code pairs make it possible to serve many tasks from one system through modular code policies .
Code policies and skill libraries
Code as Policies, announced in November 2022 by Liang and Zeng, redefined robot programming. Primitives such as pick, lift, and move-to are exposed as Python functions, and the model solves each new task by composing code from them. According to the arXiv paper 2209.07753, models produced sensible sequences on the first try without extra robot data. That one-shot success became the spark for later agent work.
Voyager carried the same idea into tool-making inside Minecraft. The agent starts from built-in Python functions, compiles new files such as francois.py during runs, and turns them into callable skills. In-context learning shines here because a few examples teach the new routine. Yet the same experiments showed this learning is non-monotonic: success wobbles first, saturates after about 20 to 40 examples, and adding more examples past a full window stops helping.
That is why the Francois ladder matters: fast in-context trials, then retrieval-backed memory, then low-rank LoRA adapters, and at the top full fine-tuning or reinforcement learning. A company with boundless driving data such as Tesla would clearly never settle for in-context methods. Compiling daytime state-action-reward traces into weights through a sleep-like pass is seen as the robotics version of the DreamCoder tradition.
Hardware demo and the latency truth
In the live demo, the agent Ashra picks a cube from a table and drops it into a bowl using only camera input. The system runs in turns: images arrive, the model outputs the next end-effector pose, and the arm moves there. At this layer code behaves less like a full policy and more like a string of tool calls . Deterministic repeats such as the approach can be compiled into code, while variable points such as object detection and failure checks keep a vision-language model in the loop.
The gap between the Waddle layer and direct Astra use shows up in economics and speed. Thinking with a giant model at every step is slow and expensive. The team therefore compiles the first successful trace into a fast skill and calls that skill on later runs. According to OpenAI, Astra reaches 64.6 on Terminal-Bench Science, ahead of the Fable 5.1 line, at roughly one-third lower cost. If the near-doubling of monthly speed continues, near real-time control could be discussed before the year ends.
Platonic representation and the next two years
Another line of support comes from cognitive science through the Platonic Representation Hypothesis. Work by Huh and colleagues at MIT shows that growing vision and language models measure distances between data points in increasingly similar ways. If that convergence holds, a strong language model already carries a strong world model. Growing world understanding through code and computer-use data then beats caging the architecture inside robot-only data.
The candidate data behind the Astra jump fits: computer use plus CAD and Blender scenes. Dragging a cursor to rotate an object teaches up-down and left-right as spatial concepts. Interfaces once modeled on the physical world at Xerox PARC now teach robots in reverse. Princeton-flavored cloud-robotics trials report that the right binding between cursor-like tools and robot targets visibly improves physical-task scores.
The closing forecast is bold: within two years, general-purpose robots that carry out any natural-language order a capable young person could do by hand. That would be a ChatGPT moment for robotics. Even the speakers warn of hard knots: latency, safety verification, and pruning a growing skill library. Still the direction is set: diversify data types, let the agent write code, keep measurement open, and compile repeats into speed.
Key moments
AI commentary
"The founders' excitement is contagious, but the RoboCurve emphasis on measurement keeps it honest. What makes this discussion valuable is that it moves the robotics debate from architecture slogans to data and engineering."
AI assessment
The strongest counterargument is latency and real-time control. Factory floors need millisecond reactions while frontier models still think in seconds per turn. Small on-device policies fill that gap today, and most deployment capital still flows there. Until the latency and cost curves bend, the puppeteer vision cannot generalize.
What is missing is a safety and liability framework. A wrongly written grasp routine can break expensive hardware or hurt a person. The conversation covers failure detection and replanning but never names a formal verification layer. Messy long-tail scenes such as wet floors, cluttered workshops, and crowded homes remain under-tested.
The speakers' incentives are visible: Waddle Labs sells agent infrastructure and RoboCurve sells credibility through measurement. Both benefit from the generalist-model wave and from an optimistic timetable. That does not invalidate the data, but the two-year general-purpose robot forecast deserves a discount.
The practical takeaway is to avoid locking robot investment to a single body. Teams that test specialist policies and cloud agents on the same arm learn fastest. For small teams the cheapest start is to give a grasp task with existing cameras to a code-writing agent and to file failures into a skill library. Those who set up measurement first will see first which jobs are truly automated.
Sources
9 links; 1 of them also cited by 17 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Y Combinator
- @web.mit.edu MIT — Robot-Use Agents
- @waddlelabs.ai Waddle Labs — Introducing Waddle
- @robocurve.org RoboCurve — Real-world evaluations
- @deepmind.google Google DeepMind — RT-2
- @research.google Google Research — Code as Policies
- @arxiv.org arXiv — Platonic Representation Hypothesis
- @openai.com OpenAI — GPT-6 Astra
Also cited by: Brain and Body: A Single-Screen Agent Setup with GPT-6 Astra on Hermes · Space Bunny Alpha: Inside OpenRouter's Free Anonymous AI Experiment · Building a Productive Card Collection App in Minutes with Base44 and GPT-6 Astra · Price War Begins: GPT-6 Sol and Luna Halve Model Costs · Gemini 4 Leak? 10 Interactive 3D Tests Against GPT-6 Astra and Fable 5.1 · Gemini 4 Pro Leaks, GPT-6 Soul in Testing: From Arena to Google Cloud, the Week's AI Shockwave · 10,000 Agents, 88 Hours, $1 Million: AI Mastermind #39 From Code to Cash to Autonomy · Cloning a Channel With One Prompt: The $33K Video Factory Built on GPT-6 Astra and Higgsfield · GPT-6 Astra Guide: How Horizontal Power Turns the Model Into Work Done · From Hand Sketch to Realistic Villa: A Showcase Video with GPT-6 Astra and Higgsfield MCP · The $500-a-Day Claim With GPT-6 Astra: Building Three Business Models End to End · When Agents Take the Job: Investing in the Stack After GPT-6 Astra (+5)
- @the-decoder.com The Decoder — Astra spatial leap
robot-use agents · llm · waddle labs · robocurve · code as policies