Back to feed

Google's Bet on Embodied AI: Gemini Robotics 2 and the Race for Physical AGI

Google's Gemini Robotics 2 push aims to give AI a body through a three-layer stack that reasons, converts thought into motion, and runs on the machine itself; yet dexterity scores and the generalization problem show the road to physical AGI is still at its start.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — L1yZ7TILj68
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

For over a decade, an invisible pane of glass stood between us and artificial intelligence: it could write, see, and solve, but it could not touch. Google is now trying to shatter that glass. The system called Gemini Robotics 2 targets embodied intelligence ; not a mind that describes the world from a distance, but one that behaves inside it.

Think of how you pick up your morning coffee: your hand reaches for the handle, fingers close, and the grip tunes itself to the weight. You compute no distances and weigh no forces against the ceramic; your brain settles years of experience below awareness. Ask a robot for the same job and the picture changes: find the cup, pick the handle, hold the angle, set the force, compensate for the weight, and keep balance throughout the carry. Every step depends on facts that may differ from the last attempt; the table can be tilted, the cup full, a person crossing. A language model can reason about a cup, but a robot must convert that reasoning into force, motion, balance, timing, and contact.

The factory robot escapes most of these troubles because the environment is arranged around the machine; it repeats the same move at the same spot thousands of times. The humanoid works in a world built for humans: door handles at human height, shelves at human reach, stairs for two legs. The speaker's thesis sharpens here: rather than single-task specialists, the generalist robot that takes on many different jobs will create more value; the running, jumping show machines are the shop window, not the real test.

A three-layer intelligence stack

On the DeepMind side, the architecture rests on three parts. At the center sits the action model: it takes camera images plus language instruction and turns them into motion. Above it, the reasoning layer called ER2 reads the scene, splits work into steps, tracks progress, and revises the plan when reality disobeys. At the bottom runs a version that lives directly on the machine, removing the round-trip delay to a distant server. As the DeepMind product pages confirm, this on-device model adapts to a new body within hours and with fewer than 200 examples. Seen together, a stack emerges: grasp the scene, decide, turn decision into motion, and execute it locally.

This stack's mind shows best in the garage-tidying scene. The user wants a messy garage cleaned and tools returned to their kit; the system grabs the scrubbing mitt and the spray bottle together and carries them to the rightmost clear bin on the top shelf. The upper model solves what to do, where to go, and the hardest question: when the job counts as done. The ability DeepMind documents as success tracking means the robot knows whether to retry or stop; in multi-robot setups, machines recognize each other's strengths and split the work.

Whole-body intelligence and the Apollo test

Google trials this mind on Apollo 2, a humanoid platform built by Apptronik. The instruction is plain: take the watering can to the green bin on the bottom shelf. Yet walking is more than swinging two legs; each step shifts the center of mass, each reach disturbs balance, each grasped object changes the forces on the body. In the Apptronik design, legs, torso, arms, and hands must work as one coordinated system; the company calls exactly this whole-body intelligence . The demo succeeds, with an honest footnote: the machine moves slower and less fluidly than a person; doing it once is a show, doing it fast and reliably every time is a product.

The real gap opens at the fingers. A human hand holds an egg without crushing it, twists a cap, ties a knot, feels the contact moment, and corrects the grip before a slip. The five-fingered hand fitted to Apollo 2 offers 22 degrees of freedom; the chance of failure grows with the flexibility. Google's evaluations put numbers on the table: unscrewing a bulb succeeds about 92% of the time, but screwing it back drops to 36%; knotting a trash bag hits 44%, sweeping into a dustpan 32%, sealing a zip bag around 40%. For a human, fitting a bulb is a near-certain job; for the robot, a dice roll. Walking may win applause, but everyday touching is the true exam of physical intelligence.

Transfer across bodies

The second test of the general-purpose claim is whether intelligence can detach from the body. Every platform differs in joints, sensors, proportions, and limits; moving a skill from an arm to a humanoid resembles performing with the right hand what was learned with a wholly different body. Google answers by running the same model checkpoint on three different bodies: Apollo 2 versions with different hands and a dual-arm rig with a gripper. The on-device model's adaptation to a new body in hours and with few examples changes the economics: if no machine needs months of special training, one intelligence can ship across bodies for different jobs. The long-term goal is a scalable intelligence layer that survives a change of hardware.

Yet a model generalizes only as far as its experience, and physical experience is costly. Language models had the internet; billions of pages could be copied digitally. The robot has no such shortcut: it cannot learn a towel's feel by reading, nor a drawer's force from a description. Four sources are counted: teleoperation where a human drives the machine remotely, videos recording people at work, simulation offering millions of repeats in virtual rooms, and data the robot gathers itself in the field. None alone covers the combinations of light, surface, object, human, and failure the physical world serves.

The data race and Robot Park

So the humanoid race turns into a race for hoarded physical experience. Figure, together with Brookfield, builds a data network spanning over 100,000 homes plus hundreds of millions of square feet of commercial and logistics space; its Helix model learns navigation purely from human head-camera recordings and finds its way in real rooms with no robot demonstration. According to Figure's reports, zero-shot transfer from human video to robot navigation has been shown for the first time. The valuable asset is no longer the robot itself; it is the pool of experience born between robots and the humans around them.

Google and Apptronik choose to manufacture experience like a factory. Per the Apptronik announcement, the Robot Park expanded in Austin on June 30, 2026 runs Apollo 2 fleets across logistics, manufacturing, and retail tasks on nearly 90,000 square feet, with customer sites such as Mercedes-Benz and GXO in the network. Every successful grasp becomes a training sample, every failed attempt another; even a passerby turns into data. From this grows the data flywheel : robots generate data, data improves models, better models enter more settings and bring back more data. Deployment becomes part of training.

Generalization and programmable work

The flywheel works only if experience converts into reusable knowledge; repeating a weak signal a million times yields no good model. A grip that sings in one warehouse can break when the light shifts; a hand that holds a dry box slips on a wet one; a move that is safe with nobody around turns risky when a human steps in. The system must learn patterns instead of memorizing, and reason about scenes it never met. Rival camps attack the same question from different sides: Google's Gemini robotics, Figure's Helix, Nvidia's GR00T platform, and the Physical Intelligence team; architectures differ, the question is one: how does physical experience turn into reusable intelligence?

The ER2 layer attempts a collective answer: split the job and deal it across robots. One stands well placed to lift, another close to the target, a third takes a different piece; the upper mind sets the division of labor. While one robot occupies a single point, a team shares space and effort, covering for a member that stalls. The difficulty is coordination under uncertainty: two robots in harmony in a controlled hall prove nothing about hundreds working alone in a chaotic depot. Contested space, broken gear, a plan collapsing midway; shared intelligence must trust reasoning, not fixed scripts, in those moments.

The economic frame shifts here too: the question is not jobs taken wholesale, but tasks becoming programmable one by one. The walking, picking, carrying, sorting, scanning, inspecting, and loading inside a shift cross the threshold separately. A machine that fits a bulb with 36% odds cannot join a production line; a machine slower than a human stays expensive in most jobs. The near term holds no millions of humanoids replacing workers, but a gradual widening of the task pool worth automating; the boundary moves outward with each gain.

Safety: knowing when not to act

When software errs, the output lands in the bin; when a physical machine errs, goods break and people get hurt. So the system's most critical skill may be knowing when to stop: telling what it sees, where its body stands, objects from humans apart, what counts as unsafe, and when to call for help. Google probes this front with the ASIMOV benchmarks; the first release, presented at CoRL 2025, generated robot constitutions automatically and lifted behavior alignment to 84.3%, published openly through GitHub. The second release tests physical-danger perception, and the agentic version scores refusing unsafe jobs, halting at critical moments, and consulting humans. The closing vision ties to this threshold: on the road to physical AGI, the question is no longer whether the machine parses the order to pick up the cup; it is whether it knows where the cup stands, how fragile it may be, who stands nearby, and what to do when the plan fails. The moment intelligence gains a body, the world itself joins the exam.

Visualization: nodesdaily AI

Key moments

  1. The behind-glass thesis
  2. The coffee cup example
  3. Three-layer architecture intro
  4. Apollo 2 watering-can demo
  5. Dexterity scores
  6. Robot Park and data flywheel
  7. The ASIMOV safety benchmark

AI commentary

"The speaker tells Google's roadmap as a convincing architecture story, but the dexterity scores reveal a machine still in apprenticeship. The front to watch is not demos; it is how fast field data turns into better models."

AI assessment

The strongest objection is that the data-flywheel story runs too smoothly. Gathering field data alone produces no generalization; teleoperation logs carry the operator's habits, and simulations carry their physics engine's assumptions. On the Figure front, navigation transfer from human video impresses, yet walking and grasping are different hardships; in contact-rich work such as holding, twisting, and feeling, the human-video signal weakens. So while data scale grows, contact richness may not keep pace, and the flywheel can slow.

The gaps in the presentation also stand out. Energy use, hourly operating cost, and how many times slower than a human the system runs never reach numbers. On the Apptronik side, Apollo 2's dimensions remain unpublished, so pasting the first-generation Apollo's figures onto the new machine would mislead. On safety, the ASIMOV benchmarks run in laboratory scenes; testing hundreds of robots at once in a crowded depot is a scale problem of another order.

The speaker's frame orbits Google and DeepMind; rival approaches get a short parade. Nvidia's GR00T platform, Figure's Helix model, and the Physical Intelligence team are named briefly and dismissed; yet their data strategies differ at the root. The Apptronik partnership and Robot Park praise also run one-sided; no independent source audits the field data's quality. None of this makes the story false, but it marks its axis: this is no industry panorama, it is a defense of Google's roadmap.

The practical takeaway for readers is sharp: humanoids will take single tasks before they take jobs. A shelf-stocking, box-carrying, or simple tidying job automates once it crosses the reliability threshold; no investment decision should precede that crossing. Three metrics deserve tracking: success rates across repeated trials, the cost of fitting a new body, and how fast field data converts into model updates. Any pilot bought before these three improve stays an expensive show.

Sources

6 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

artificial intelligence · robotics · google deepmind · humanoid robot · automation

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…