For years, robots have been stuck in a frustrating middle ground. Industrial arms excel at repeating the same motion millions of times, but they struggle when something unexpected happens. Research robots can perform impressive tricks in tightly controlled labs, yet they often fail when asked to simply walk across a room and pick up an object that was slightly moved. Google's DeepMind team has been working to change that with a new generation of AI models that give robots a far more human-like sense of balance, dexterity, and coordination.
The latest breakthrough, Gemini Robotics 2, builds on the company's original Gemini Robotics model and takes a major leap forward. Instead of focusing on isolated tasks like stacking blocks or picking up objects from a table, this new system is designed to control an entire humanoid robot from head to toe. The result is a machine that can walk, reach, grasp, manipulate, and even collaborate with other robots in real time. It marks a significant step toward robots that can operate safely and usefully in real-world environments, where the unexpected is the norm.
From tabletop tasks to whole-body control
One of the biggest limitations of earlier robotic AI models was their narrow scope. Most were designed for tabletop work, where a robot sits in one spot and uses only its arms, wrists, and grippers to interact with objects within reach. That was useful for research and for a handful of industrial applications, but it did not reflect the demands of real-world tasks. Humans rarely stay in one place; we walk across rooms, bend down, reach up, and adjust our posture constantly. For robots to truly help in homes, warehouses, and hospitals, they need that same whole-body awareness.
Gemini Robotics 2 changes that by taking control of the entire humanoid form, from the feet that maintain balance to the fingertips that perform fine manipulations. In a demonstration with Apptronik's Apollo 2 robot, the AI was given a simple instruction: place a watering can into a bin on a bottom shelf. The robot did not need a pre-programmed route or a teleoperator guiding its movements. It walked over to the watering can, bent down, picked it up, crossed the room, and set it down precisely where it belonged. What makes this remarkable is that the robot had to coordinate its walking gait, its balance, its arm trajectory, and its grip strength all at once, while adjusting to the exact position and orientation of the bin.
This kind of whole-body control is a prerequisite for any meaningful robot assistant. It allows a machine to navigate cluttered environments, avoid obstacles, and complete multi-step tasks without constant human intervention. The system understands that a robot's center of mass shifts as it moves, and it takes that into account with every step. The result is a smoother, more natural motion that closely mimics how a person would approach the same task.
Hands that can tie knots and seal bags
Dexterity has always been one of the hardest challenges in robotics. Human hands have more than 20 degrees of freedom, allowing us to twist, pinch, grip, and manipulate objects with incredible subtlety. Robots, by contrast, often rely on simple two-fingered grippers that can only perform basic pick-and-place actions. Even multi-fingered hands have historically been slow and clumsy, with each finger controlled by a separate script that struggles to adapt to new objects.
Gemini Robotics 2 brings a dramatic improvement in this area. The model can control a five-fingered robotic hand well enough to tie a knot, seal a zip-lock bag, or unscrew a light bulb. These are tasks that require precise force control, tactile feedback, and the ability to adjust on the fly. Tying a knot, for example, is not just about moving a string; it requires sensing the tension in the material and coordinating multiple fingers simultaneously. The fact that an AI model can achieve this is a strong sign that robots are moving beyond rigid automation and into flexible manipulation.
At the same time, the model works just as smoothly with simpler two-fingered grippers. For tasks like packing and sorting, a parallel jaw gripper is often more reliable and faster than a human-like hand. Gemini Robotics 2 can switch between these different end effectors without any loss of performance. That versatility makes it useful across a wide range of industries, from e-commerce fulfillment centers to food processing facilities, where different tasks require different levels of dexterity.
A reasoning brain for long-horizon tasks
Physical dexterity is only part of the equation. A robot that can tie a knot is impressive, but it is far more valuable if it can also understand what task it is supposed to do, how to sequence the steps, and what to do when something goes wrong. To address this, Google also introduced Gemini Robotics ER 2, a reasoning model that acts as the robot's project manager.
This reasoning model is designed to break down high-level instructions into actionable steps. If you tell a robot to clean up a spill and put away the supplies, it will split that into sub-tasks: locate the spill, find a mop or towel, clean the area, return the cleaning supplies to their storage spot, and then report back. The model keeps track of multi-minute tasks, which is a significant departure from earlier systems that lost focus after a few seconds. It can remember what has already been done, what still needs to be done, and how to adjust the plan when new information comes in.
One of the most exciting capabilities is multi-robot coordination. Gemini Robotics ER 2 can manage several robots working on the same job, assigning different roles to each one and ensuring they do not get in each other's way. Imagine a warehouse where one robot fetches boxes from shelves, another carries them to a packing station, and a third wraps and labels them. The reasoning model coordinates all of this, acting like a conductor leading an orchestra of machines.
There is also an on-device version designed for robots that do not have reliable internet access. The on-device model can adapt to a brand-new robot body in just a few hours, using as few as 200 examples of the new robot's movements. This is crucial for real-world deployment, where manufacturers rarely want to rely on a cloud connection for every robot in a factory. It also speeds up the integration process, allowing new robots to become operational much faster than traditional machine learning pipelines, which typically require thousands or even millions of training examples.
Safety at the center
As robots become more capable, safety becomes even more important. A robot that can tie a knot or walk across a room also has the potential to cause harm if it misjudges its surroundings. Google has made significant efforts to address these concerns with Gemini Robotics ER 2, which includes multiple layers of safety.
The company introduced a new benchmark called ASIMOV-Agentic to test whether robots can make safe decisions in a variety of scenarios. The benchmark evaluates whether a robot knows when to refuse a risky action or ask a human for help instead of blindly following an instruction. This is a subtle but essential capability. For example, if a robot is told to move a heavy object but its sensors detect that the object is too heavy and likely to fall, the best course of action is not to attempt the lift. It should either refuse and explain why, or find a safer alternative approach.
Gemini Robotics ER 2 also includes proximity awareness. The model can sense when a person gets too close to the robot and bring the robot to a safe stop. This is particularly important in environments where humans and robots share the same space, such as factories, hospitals, and homes. A robot that freezes when a person enters its workspace is much safer than one that continues moving without regard for the person's presence. This kind of behavior is a step toward the collaborative robots that safety regulations envision, where machines and people work side by side without the need for physical barriers.
The focus on safety is not just about avoiding accidents. It is also about building trust with users and regulators. For robots to be accepted in daily life, people need to feel confident that they will not be harmed. By embedding safety into the core reasoning process, Google is making it clear that responsible robot behavior is a priority, not an afterthought.
Availability and what's next
Gemini Robotics ER 2 is already live on Google AI Studio, which allows developers and researchers to experiment with the reasoning model and understand its capabilities. The rest of the models, including the full Gemini Robotics 2 system, are currently rolling out to early access partners. These partners will be the first to deploy the technology in real-world applications, providing critical feedback to Google about performance, limitations, and potential improvements.
The introduction of Gemini Robotics 2 is part of a broader trend in the field of embodied AI. Companies like Tesla, Figure, and Boston Dynamics have all been working on humanoid robots that can navigate everyday environments. What has been missing, until recently, is a robust software layer that can make these physical machines genuinely intelligent. With Gemini Robotics 2, Google is positioning itself as a leader in that software layer, offering a model that other robot manufacturers can integrate into their hardware.
The implications are enormous. Humanoid robots equipped with this kind of AI could eventually perform household chores, assist in elder care, and handle dangerous jobs in disaster zones. They could also transform industries like logistics and manufacturing, where there is a constant need for flexible automation. However, there are still significant hurdles to overcome. Battery life, hardware cost, and physical robustness remain major limitations. And while the AI improvements are impressive, they have not yet been tested at scale in the chaotic, unstructured environments of the real world.
What is particularly promising is the speed of progress. The original Gemini Robotics was announced only months earlier, and already the successor is demonstrating capabilities that seemed like science fiction a decade ago. The ability to adapt to a new robot body in a few hours with minimal examples is a game-changer for the robotics industry. It means that advances in AI can be rapidly transferred to any new hardware that comes to market, accelerating the pace of innovation across the entire sector.
Gemini Robotics 2 is not just a step forward for Google; it is a step forward for the entire field of robotics. It brings together the three pillars of intelligent action: perception, reasoning, and motor control. By integrating these into a single model, Google has created a blueprint for what useful, safe, and adaptable robots will look like in the years to come. As early access partners begin to test the limits of this technology, the world will get its first real glimpse of a future where robots are not just tools that follow commands, but autonomous assistants that understand their surroundings and work alongside us with remarkable skill.
Source: Digital Trends News