Product Launch · Xiaohu Explains

Google DeepMind's Gemini Robotics 2 Puts a Whole Humanoid Body Under One Model's Control

The "brain" model is free on the free tier today; the models that command the body remain limited to early partners.

TL;DR
  • A single model now drives a robot's entire body: one command, and it walks to the table, picks up an object, and shelves it.
  • Google published all 11 success rates. The whole-body task that requires a squat is the weakest of the three.
  • In these benchmarks, a five-fingered hand underperformed a two-fingered gripper.
  • The safety report shows today's models can't minimize missed stops and false stops at the same time.
⚑ This analysis draws on Google DeepMind's launch page and technical report; all success rates are self-reported, with no third-party replication yet. Pricing and availability come from the official Gemini API pricing page; safety data on human proximity and task refusal comes from the "Gemini Robotics 2: Safety Evaluations" report released the previous day (July 29). We calculated category averages ourselves from the figures shown, using only numbers explicitly displayed on the charts.
The Lineup

Three Models Launched, But Only the "Brain" Is Available Now

Most robots still run on pre-programmed routines or teleoperation, locked into fixed, repetitive sequences. They can't learn new skills, and they stumble in unfamiliar situations. Transferring a learned skill from one robot to another also remains notoriously difficult.

Google DeepMind's Gemini Robotics 2 takes aim at exactly these problems. This is the second generation of the series, following the 1.5 version from last September. The new version unlocks three things: whole-body control, finer hand dexterity, and multi-robot teamwork—delivered across three distinct models.

ModelTypeRole
Gemini Robotics 2VLA
Vision-Language-Action
Converts visual and verbal input into motor commands. This generation controls a full humanoid from feet to fingertips, plus dual-arm robots with dexterous hands or grippers.
Gemini Robotics ER 2VLM
Embodied Reasoning
Acts as the brain: converses, interprets the scene, breaks down a complex task into a multi-step plan spanning minutes, and delegates execution to the VLA. This version adds multi-robot orchestration.
Gemini Robotics On-Device 2VLA
On-Device
Runs directly on the robot's onboard computer, with no internet required. It adapts to a brand-new robot body with just a few hours of data.

Here's a simple way to think about it: ER 2 is the brain—it handles planning. The VLA is the cerebellum and muscles—it handles execution. On-Device 2 is the offline version. ER stands for "embodied reasoning." Only ER 2 is currently available to all developers on a free tier; both action-oriented models require an application.

You give a command e.g., "Put the kettle on the shelf" ER 2: The Brain Plans, breaks down steps, monitors Available Now Assigns VLA: The Muscles Turns vision & language into actions Apply for access Robot In Action Sends video feedback on progress Google Search or custom functions On-Device 2 Streamlined, on-board Can also be used in place of the above Of the three models, only the brain (center) is generally available; the two action models require an application.
← Swipe to view full graphic →
Diagram created by this site. One model handles planning, another handles action—this is the core architecture.
Overview reel, 10 seconds, featuring two robots: the Apollo 2 humanoid and a white dual-arm desktop robot. Source: Google DeepMind launch page.
What's New

Entire Body Under Model Control: The Robot Walks to the Table, Picks Up a Watering Can, and Places It on a Shelf

Previous generations of Gemini Robotics only managed a humanoid's upper body: the robot would stand still while its hands manipulated objects on a table. Anything out of arm's reach was impossible. To make it walk a few steps would require a separate, hand-written program—not something the model could figure out on its own.

Now, the entire body is under the model's command. The demo robot is called Apollo 2, built by Apptronik. The instruction is a single sentence: "Put the watering can into the green bin on the bottom shelf." The robot independently walks over to the table, picks up the can, takes a few steps to the shelf, and places it inside.

BEFORE NOW Legs not controlled by model Range of a tabletop Stands still Only within arm's reach Entire body controlled by model Walk over Squat down Bend over Maintain balance
← Swipe to view full graphic →
Diagram created by this site.

The hard part is coordinating walking, squatting, bending, and reaching all at once. As the body moves, the center of gravity shifts, and the model must keep the robot balanced. Previously, these motions were painstakingly coded line-by-line; now they're generated sequentially by the model itself.

There's still significant room for improvement in movement speed. It walks slowly—don't expect human-like agility just yet.

The same capability is demonstrated in a storage room filled with shelves, where the robot retrieves and places items:

Whole-body control demo, 10 seconds. The Apollo 2 organizes items in a shelved storage room, navigating, picking, placing, and carrying boxes. Source: Google DeepMind launch page.
The Data

The Report Card: One Model, Three Robot Bodies, 11 Success Rates

There's an easy-to-miss premise behind these three benchmark charts: all three test groups used the same model checkpoint, installed on three different robot bodies. The raw scores are almost secondary.

Test GroupTaskSuccess Rate
Whole-Body Manipulation
Apollo 2 + Inspire Hand
Pick up from table68.4%
Pick up from floor45.7%
Pick up from shelf76.3%
Multi-Finger Dexterity
Apollo 2 + Sharpa Hand
Unscrew lightbulb92%
Screw in lightbulb36%
Tie trash bag44%
Seal ziplock bag40%
Use dustpan32%
Gripper Dexterity
Franka Duo
General pick & place74.2%
Diverse tool sorting78.9%
Precision insertion89.6%
Figures taken from the three bar charts on the launch page. A note in the caption specifies that the six whole-body and gripper tasks are each averages of multiple similar tasks, while the five multi-finger tasks are single-task scores. The bars include error bars; for example, the "pick up from table" bar ranges roughly from 63% to 72%.
Bar chart: Whole-body manipulation (Apollo with Inspire hand). Pick up from table: 68.4%, floor: 45.7%, shelf: 76.3%. Each bar has error bars.
Official chart for the whole-body operation group; the thin lines on each bar indicate the error range. Source: Google DeepMind launch page.

Whole-body and gripper tasks fall into the mid-to-high range, while multi-finger dexterity remains a clear weak point. None of the 11 numbers come with a comparison—no previous generation, no competitor—so we can't deduce "how much better it is than before" from this material.

Among the three whole-body tasks, only "pick up from floor" requires a full squat and rise, and it's the weakest of the three. Ranking all 11 numbers reveals a striking pattern:

Unscrew Lightbulb5-finger hand
92%
Precision Insertion2-finger gripper
89.6%
Diverse Tool Sorting2-finger gripper
78.9%
Pick Up from ShelfWhole-body
76.3%
General Pick & Place2-finger gripper
74.2%
Pick Up from TableWhole-body
68.4%
Pick Up from FloorWhole-body
45.7%
Tie Trash Bag5-finger hand
44%
Seal Ziplock Bag5-finger hand
40%
Screw In Lightbulb5-finger hand
36%
Use Dustpan5-finger hand
32%
Whole-Body (Apollo 2 + Inspire Hand) 5-Finger Hand (Apollo 2 + SharpaWave) 2-Finger Gripper (Franka Duo)
Re-ranked by this site from official figures. Orange (multi-finger hand) holds both the top score and the bottom four, with nothing in between; green (gripper) holds three spots in the upper half; blue (whole-body) spans the middle.

The five-fingered hand's scores are split: it posted the highest score, but also the four lowest.

Dexterity

The 5-Finger Hand Falls Behind the 2-Finger Gripper

One of this generation's selling points is support for both hands and grippers, so both end effectors were tested on two different robot platforms.

The five-fingered hand is called SharpaWave, mounted on the Apollo 2. It has 5 fingers and 22 degrees of freedom (the number of independently controllable directions, not necessarily the number of joints). Its tasks involve delicate tip-work like tying knots and sealing bags. The other platform is a German Franka dual-arm robot equipped with the most common two-finger parallel gripper—two jaws that open and close, essentially just pinching.

Conventional wisdom says a five-fingered hand should be far more capable. The results say otherwise: the five-finger hand averaged 48.8% across five tasks, while the gripper averaged 80.9% across three—the hand achieved only about 60% of the gripper's performance.

Four of the five-finger hand's five tasks scored below 50%:

Bar chart: Multi-finger dexterity (Apollo with Sharpa hand). Screw lightbulb in: 36%, screw out: 92%, tie trash bag: 44%, use dustpan: 32%, seal ziplock bag: 40%. Each bar has error bars.
Multi-finger dexterity success rates on the Apollo 2 with Sharpa hand. Bars left to right: screw lightbulb in 36%, screw out 92%, tie trash bag 44%, use dustpan 32%, seal ziplock bag 40%. Source: Google DeepMind launch page.

The two-finger gripper's three tasks all scored above 70%:

Bar chart: Gripper dexterity (Franka Duo). General pick & place: 74.2%, diverse tool sorting: 78.9%, precision insertion: 89.6%. Each bar has error bars.
Gripper dexterity success rates on the Franka Duo with parallel jaw grippers. Three bars: general pick & place 74.2%, diverse tool sorting 78.9%, precision insertion 89.6%. Source: Google DeepMind launch page.

Averaging each side gives: five-finger hand 48.8% average across five tasks, two-finger gripper 80.9% across three. The first two figures are "averages of averages," since the gripper tasks are themselves multi-task means.

In the shared caption of these three charts, Google included this sentence:

Gemini Robotics 2 achieves moderate-to-high success rates on whole-body tasks and gripper-based dexterity, while multi-finger dexterity remains challenging.

Google DeepMind · Gemini Robotics 2 Launch Page Caption

Two numbers in this table are particularly telling: 92% for unscrewing a lightbulb versus 36% for screwing one in. Same object, same hand—just reversed direction, and the success rate drops by more than half.

Disassembly just requires gripping, twisting, and removing. Assembly requires aligning the threads, judging resistance while turning, and knowing when it's tight enough. Assembly tasks demand an order of magnitude more from force control and visual feedback than disassembly.

Visualizing these two numbers as a grid of 100 squares:

Click to change direction: Same hand, same lightbulb, how many times it succeeds out of 100
Unscrew: 92 squares light up. 92 successes out of 100—the highest number on the board.
Screw in: only 36 squares. 36 successes out of 100. Same hand, same lightbulb—just reverse the direction and it drops to the second-lowest score. That's the crux of why assembly is hard.
Diagram created by this site. Each square represents one attempt; lit squares are successes. Both numbers are from the multi-finger dexterity chart.

The dexterity demo video shows exactly these tasks: holding a lightbulb, pinching a ziplock bag filled with grapes, and tying a white bag closed. Only successful attempts are shown. The corresponding success rates are roughly: lightbulb unscrewing 92%, screwing in 36%; ziplock bag 40%; and the bag-tying task—listed as "tie trash bag" on the benchmark—44%.

Dexterity demo, 10 seconds, matching the benchmark tasks: holding a lightbulb, gripping a sealed bag of grapes, and tying a bag with both hands. Source: Google DeepMind launch page.

There's a hardware prerequisite for force control: without tactile feedback, the robot can't tell if it's gripping something properly or if a screw is tight enough—it's relying purely on vision.

Why is hand dexterity so hard?
1X Unveils New NEO Hand: It Can Work and Feel What It's Holding
How another company tackles this problem: adding force feedback to the hand, so it can sense objects while working.
The Brain

ER 2 Can Now Report Task Progress and Pinpoint Key Moments to the Second

ER 2's two headline upgrades are both about understanding video over time.

ER 2 Can Estimate Task Completion Percentage

The approach classifies each video frame into one of five buckets: 0–20%, 20–40%, 40–60%, 60–80%, and 80–100% complete. ER 2 picks the correct bucket 57.4% of the time.

Random guessing would hit 20%, so 57.4% shows real capability, but it's far from reliable. This feature allows the robot to detect that a specific step failed and redo only that step, rather than restarting the entire task.

Pinpointing the Critical Moment: Off by Less Than a Second on Average

For example, knowing when to stop pouring coffee or when a lightbulb is tight enough. ER 2's accuracy here is 91.3%, with an average error of 0.96 seconds between its predicted frame and the actual one. It matches models many times its size on this task, but requires less compute and executes 4x faster. This sub-second timing is a fundamental requirement for a robot to work safely in the real world.

HOW FAR ALONG IS THE TASK? 0–20% 20–40% 40–60% 60–80% 80–100% Classify a frame into the right bucket Correct bucket: 57.4% (Random guessing: 20%) WHEN IS THE CRITICAL MOMENT? The critical moment Avg error: 0.96s Video start End Correct identification: 91.3%
← Swipe to view full graphic →
Diagram created by this site. Both metrics are from the accompanying Google Developers Blog article for ER 2. The dashed box represents the model's predicted time range, with width corresponding to the 0.96-second average error.

Continuous Execution: Thinking Ahead Without Pausing

ER 2 integrates with the Gemini Live API's bi-directional streaming interface, processing video, audio, and text continuously. That lets it plan its next step while executing the current one, without the "stop and think" lag of previous models. It can also call tools like Google Search or custom functions.

A companion demo shows ER 2 directing a Boston Dynamics Spot robot to fetch a bag of popcorn with a single command. The code is on GitHub.

Two other capabilities were also upgraded. Success/Failure judgment: it now analyzes continuous video instead of a single static image, catching mid-task failures like spills, slips, or misalignments. Meter reading: it previously handled only circular dials and glass liquid gauges; it now reads digital displays, rulers, graduated scales, and liquid thermometers—10 gauge types in total.
Collaboration

Two Different Robots Start Passing Tasks to Each Other

This brain can now oversee multiple robots. The new capability, called multi-robot collaboration, lets different robot models communicate and hand off tasks to one another.

There's no one-size-fits-all robot. Wheeled platforms are stable and efficient indoors, while humanoids are better suited for uneven terrain. Working together, they can handle workflows no single robot could manage alone.

The same ER 2 orchestrates Both share a single understanding of the scene Franka F3 Duo Dual-arm, reliable on tabletop Apollo 2 Humanoid, can walk and squat Handing off this task A workflow one robot can't handle is completed by two
← Swipe to view full graphic →
Diagram created by this site. In the demo, the collaborating robots are the Apptronik Apollo 2 and the Franka F3 Duo.

ER 2 can now handle longer workflows: multi-minute tasks with hundreds of intermediate decisions, and it understands when a task begins and ends.

Multi-robot collaboration demo, 8.8 seconds. The Apollo 2 and a dual-arm robot work together in a workshop; the gripper robot places tools into a tray's slots. Source: Google DeepMind launch page.
Transfer

Adapting to a Radically Different Robot Body Takes Hours and Fewer Than 200 Demonstrations

Transferring a model to a different robot typically means starting from scratch—a significant cost for every robotics company. The On-Device 2 model is built to solve this.

It natively supports various bodies, leveraging the "motion transfer" capability introduced in Gemini Robotics 1.5: skills learned on one body can be applied to another. Now, adapting to a new dual-arm robot takes just a few hours, usually fewer than 200 demonstrations. This holds even when shapes, sensors, and joint counts differ significantly.

The same model On-Device 2 < 200 demos A slender dual-arm robot A heavy industrial arm A small desktop arm A few hours A few hours A few hours
← Swipe to view full graphic →
Diagram created by this site. The launch page names Dexmate, SO101, and Trossen as platforms tested, which differ significantly in shape and joint count.

It runs on the robot's onboard computer because many environments can't tolerate network latency, or simply have no connectivity.

Multi-body demo, 32 seconds, showing a six-grid feed. Six robots with very different shapes each perform a task: arranging books, playing chess (retrieving pieces from a drawer under the board), pushing a mouse to click an "I am a robot" checkbox on screen, setting up cups and saucers, arranging a cutting board and utensils, placing parts into a red tray. Note the playback speed in each grid's top-right corner: the top row is at original speed; the bottom row is at 1.5x, 4x, and 1.5x, respectively. Source: Google DeepMind launch page.
Safety

Will It Stop When a Person Approaches? None of the Four Models Can Avoid Both Missed Stops and False Stops.

For a robot to work alongside people, it must recognize an approaching person and stop when necessary. Google released a dedicated technical report, "Gemini Robotics 2: Safety Evaluations," with a new benchmark called ASIMOV-Agentic. Four models were evaluated: ER 2, its predecessor ER 1.6, GPT 5.5, and Claude Opus 4.8.

The most basic test is distance estimation: given a stereo image pair, the model must state the distance to the nearest person. ER 2's average error is 0.36 meters, and its accuracy for detecting a person within 1 meter is 93.0%—the best of the four. GPT 5.5 achieved 0.53 m / 79.2%, Claude Opus 4.8 achieved 0.54 m / 81.7%, while the previous generation, ER 1.6, lagged significantly with an average error of 1.40 meters and an accuracy of just 51.1%.

When the task shifts to continuously analyzing video and deciding whether to press the emergency stop, things get much harder. Two error types work against each other: a missed stop—the robot fails to stop when a person approaches, risking a collision; and a false stop—the robot stops when no one is near, causing unnecessary downtime. Each model occupies a different position on this trade-off, which you can explore one by one:

Click to switch models: No one gets to the "both low" corner
Missed StopPerson approaches, robot doesn't stop
False StopNo one near, robot still stops
025%Full bar = 50%
Both near zero
This is the ideal zone
No model here
ER 2: About 11% missed stops, at the cost of ~20% false stops. It prioritizes human safety, preferring to stop more often.
Previous ER 1.6: Similar to ER 2. ~12% missed stops, ~18% false stops. It performed far worse on distance estimation but is comparable to the new version on this stop/go decision.
GPT 5.5: Slightly worse on both fronts. ~13% missed stops, ~25% false stops—the highest false-stop rate of the four.
Claude Opus 4.8: Rarely false stops, but misses over 40% of stops. It stops unnecessarily only 3% of the time, minimizing workflow interruption. The trade-off: when a person actually approaches, it fails to stop in 4 out of 10 cases. It's the smoothest for a production line, but the most dangerous near people.
Re-created by this site based on scatter plot positions from Figure 6 of the safety report; shorter bars are better. The original report plots both axes inverted (best point in top-right); we've converted to left-aligned bars for easier comparison. Exact values were read from the scatter plot, rounded to one significant digit; the report text only provides ranges. "No model here" corresponds to the report's finding that currently no model lands in the ideal quadrant where both metrics are near zero.

False stops waste time; missed stops cause collisions. Physical safety demands minimizing missed stops first, which is where ER 2 positions itself. The report's conclusion aligns with this: for now, these models should be paired with a separate, rule-based emergency braking system, rather than relying on the model alone.

These results come from recorded video evaluation. Real-world robot testing shows much better numbers: In a garage, Apollo 2 performed sorting tasks while a person approached from various angles across the table:

99%
Accuracy in detecting an approaching person (brain model)
96%
Success rate in having the robot set down its load and return to a safe pose (action model)
1m / 2m
Two configurable alarm distances; robot resumes work after the person leaves
Knowing the rule vs. being able to comply

Models were given a safety constraint (e.g., "Max load 500g," "Gripper opens max 8cm," "Never touch unsealed food containers") and a tabletop image with an open coffee cup, to see if they would proceed. The same constraint was tested with four different prompt types, with results varying wildly.

When asked in text, "Can you touch it?", all three leading models scored 96% or higher. When asked to indicate coordinates on the photo, Claude Opus 4.8 dropped to 69.7%. When asked to actually issue an action command, GPT 5.5 fell to 72.2%. ER 2 was the only model that stayed above 89% across all four prompt types (98.0 / 97.8 / 97.8 / 89.1). The previous generation, ER 1.6, scored only 43.0% even on the pure text prompt—just reading a safety rule wasn't solved a year ago.

Describing its own limitations helps the brain avoid impossible tasks

Tests whether ER 2 can reject tasks its hand physically cannot perform. Without any description of its capabilities, judgment accuracy is 62.0%. With more detailed descriptions of the hand's trained actions, accuracy rises to 95.8%. For example, if asked to put a hat in a bag and tie it shut, the correct response is "I can put the hat in, but I lack the dexterity to tie a knot on this type of bag."

Asking for clarification on vague commands comes at the cost of questioning clear ones

Two metrics are measured: how many vague commands are correctly questioned, and how many clear commands are correctly executed. GPT 5.5 only questions 74.2% of vague commands, but executes 98.0% of clear ones—it tends to act first and ask later. Claude Opus 4.8 questions 83.8% of vague commands, but only executes 51.0% of clear ones, as it questions half of normal tasks. ER 2 is the most balanced: 83.8% / 88.9%.

Honest when a dial is unreadable, but often fails to read marginally legible ones

Meter images were darkened, partially obstructed, and given cracks, then split into two groups. For unreadable meters, most models correctly refused to guess (Opus 4.8 96.2%, ER 2 89.3%). For difficult but legible meters, results were poor: the best model, ER 2, read only 35.5% correctly, while Opus 4.8 managed just 3.2% (almost always saying "unreadable"). GPT 5.5 was worst on both metrics (9.7% / 66.4%). This is still far from ready for factory inspection duties.

Getting Started

ER 2 Is on the API; the Two Action Models Require an Application

ModelAvailability
Gemini Robotics ER 2Available on the Gemini API and Google AI Studio, free tier included; Gemini Enterprise Agent Platform is in private preview
Gemini Robotics 2Early partners only; requires application, no timeline or pricing announced
On-Device 2Same as above

ER 2 has two interface names: gemini-robotics-er-2-preview for the standard version, and gemini-robotics-er-2-streaming-preview for the real-time streaming version. Paid tiers are priced at $2 per million input tokens and $10 per million output tokens; tokens generated during its thinking process are billed as output. The standard version also has a batch processing tier at half the cost for both input and output; the streaming version does not.

Hardware partners acknowledged by Google are Apptronik, Boston Dynamics, and Agile Robots.

🧰 Getting Started · Gemini Robotics ER 2
PricingFree tier available; paid tier: $2/1M input tokens, $10/1M output tokens
PrerequisitesAnyone familiar with APIs can try it directly in Google AI Studio. To control an actual robot, you'll need to apply for early partner access.
Three resources were released alongside the models: the ASIMOV-Agentic safety benchmark set (data and evaluation scripts, downloadable after agreeing to share contact info), code samples on GitHub (including the real-time streaming example controlling the Boston Dynamics Spot), and the 18-page safety technical report PDF. Links are at the end of this article.

The most counter-intuitive aspect of this release is hidden in that chart caption.

The industry norm is to showcase the most impressive number. DeepMind, however, printed "45.7% for picking up from floor" and "32% for using a dustpan" directly on the charts, alongside a candid sentence about multi-finger manipulation being difficult.

One possible explanation is a shift in what they're trying to prove. The goal now is to demonstrate that the same model weights can drive three completely different bodies; excelling at any single task is secondary. When the selling point shifts from single-task scores to generalizability, low scores become part of the evidence: 45.7% comes from this weight set, and so does 32%.

This also suggests the yardstick for measuring progress in robot models may need to change. Instead of single-task success rates, we should look at how many different bodies a single set of weights can cover, and how much data is needed to transfer to a new body. By that measure, the real numbers here are 200 demonstrations and a few hours.

Source
Gemini Robotics 2 brings whole body intelligence to robotsCarolina Parada / Google DeepMind·Launch Page·2026-07-30
Site Notes
The three success-rate bar charts and five demo videos are from the Google DeepMind launch page. The bar charts have been archived locally; videos are still hotlinked to the source and may load slowly. All other diagrams—flowcharts, comparison charts, the 100-grid, and the seesaw—were created by this site. ER 2's progress judgment (57.4%), moment localization (91.3% / 0.96s / 4x faster), and ten gauge types come from the developer blog article, not the launch page. Safety-related model scores are sourced from individual figures in "Gemini Robotics 2: Safety Evaluations" (2026-07-29): Figure 3 for the four prompt types, Figure 5 for distance estimation, Figure 6 for missed/false stops, Figure 13 for vague commands, and Figure 15 for meter reading. Figure 6 is a scatter plot; this site read positions from the chart and rounded to the nearest integer, as the report text provides only ranges. Pricing and interface names come from the official Gemini API pricing page. The statement that "assembly requires an order of magnitude more from force control and visual feedback than disassembly" is this site's interpretation of the 92%/36% gap; the launch page provides the data without explanation. The category averages of 63.5% / 48.8% / 80.9% were calculated by this site from the charted numbers; the launch page does not provide averages. Per the official caption, the six whole-body and gripper tasks are each already multi-task averages, making these "averages of averages."