AI 日报hiw3c.com

机器人领域的人工智能突破不会很快改变你的生活

原文标题 · AI breakthroughs in robotics won’t change your life any time soon
MIT Tech Review AI www.technologyreview.com RSS 全文
正文为英文,可一键机器翻译(仅首次需要等待)

The story is a collaboration between MIT Technology Review and Aventine, a non-profit research foundation that creates and supports content about how technology and science are changing the way we live.

A robot shaped like a human—white with a black head and torso—has been popping up on video feeds. Perhaps you’ve seen it dance or pass popcorn, put trash in a bin, vacuum, or press the button of a microwave. Or maybe you’ve watched it fall backward while handing out water bottles or struggle to iron a shirt. 

This would be Tesla’s Optimus, an AI-powered humanoid robot that Elon Musk, the company’s CEO, believes will be “not just Tesla’s biggest product ever, but probably the biggest product ever,” headed to work on factory floors and, later, in our homes. Eventually it “will have human and then superhuman dexterity,” he told shareholders in July. Optimus robots could automate almost all human labor—from hauling sheet metal to folding laundry—for as little as $20,000 each, Musk argues. Speaking at the World Economic Forum’s annual meeting in Davos, Switzerland, in January, he predicted they could be on sale to the public by the end of 2027.

Musk is not alone in his evangelism. Marc Andreessen, cofounder and general partner of the Silicon Valley venture capital firm Andreessen Horowitz, has said that robotics could become the “biggest industry in the history of the planet.” In January, Jensen Huang, CEO of Nvidia, said that humanoid robots would match human-level ability this year. According to Morgan Stanley, the number of robots that “resemble and act like humans” is likely to reach nearly 1 billion by 2050, creating a market worth over $5 trillion. 

Elon Musk predicts that Tesla’s Optimus humanoid could be one the world’s best-selling products. For now, it’s most often seen handing out food and drinks at Tesla events.
SIPA VIA AP IMAGES

Such proclamations are in large part fueled by the idea that the same AI advances behind tools like OpenAI’s ChatGPT and Anthropic’s Claude will enable a new generation of robots to imitate human movement the way chatbots imitate human language. But many robotics researchers are skeptical, arguing that such assumptions minimize the challenges of using an intelligence built on language and images to master the infinite variability of the physical world. “None of those companies [building humanoid robots]—absolutely none of them—has any idea how to make those robots smart enough to be useful,” Yann LeCun, often referred to as one of the godfathers of AI, said at another event during the January Davos conference.  

Researchers also point out that the tendency to conflate humanoid robots made to resemble people with so-called generalist machines able to learn and perform multiple tasks is misleading. ”It’s very easy to make a robot that looks like a person,” explains Jonathan Hurst, cofounder and chief robot officer of Agility Robotics and professor of robotics at Oregon State University. “It is dramatically more difficult to make a machine that moves or behaves dynamically or physically like a person.”  

These tensions—over whether all-purpose humanoid robots are just around the corner or nowhere in sight, and whether current forms of AI are all that’s needed to perfect them—are playing out in robotics labs across the country, where the hype over timelines is obscuring painstaking but meaningful progress.

A decade or so ago, a series of breakthroughs led to a generative AI revolution that turned the long-imagined possibility of artificial intelligence into reality. Roboticists—though they disagree on exactly when this will happen—believe that an equally transformative revolution is possible in robotics, one that will endow machines with physical intuition and fluidity that has long been out of reach. As progress in robotics inches forward, the question is whether the same methods and tools that fueled advances in AI are enough to get there, or if an entirely new path is required.

Robots meet advanced AI

To see one of the smartest robot brains working today, it’s worth looking at what Google DeepMind can do with a piece of equipment called ALOHA 2, short for “A Low-cost Open-source Hardware System for Bimanual Teleoperation.” 

Roboticists have long clashed over whether a humanlike form is necessary for generalist robots, with proponents arguing that it will help them slot into the world as it exists and detractors saying it’s not worth the trouble. ALOHA 2 reflects this second way of thinking. Not much to look at, it’s just a pair of arms, some grippers, and a couple of cameras. But despite its seeming simplicity, it is a workhorse for researchers at Google DeepMind, who use it to test their most advanced AI for robotics system, Gemini Robotics, in their various labs.

When controlled by Gemini Robotics, ALOHA 2 becomes more of a generalist robot, in the sense that it can perform any number of tasks based on examples it’s been trained on. Ask it to pack a lunchbox and, as evidenced by a video of this exercise, it can use two pincer grippers to delicately place a piece of white bread into a Ziploc bag, close it, place a bunch of grapes in a Tupperware container, secure the lid, and then carefully move the items into a lunchbox before zipping it up. 

It’s not a great lunch. But the fact that the robot can put it together represents an objective step forward from what was possible even, say, three years ago. 

This is in large part due to AI and its impact on what are known as robot policies, which controls how a general-purpose robot will need to assess and understand its surroundings, plan how to move within them, and then perform its task correctly.

The ALOHA 2 robot isn’t much more than two mechanical arms on a bench top, but it serves as a testbed for cutting-edge AI robotics models. These animations are based on human teleoperation of the robot arms, data that is used to train Google DeepMind’s models. (Video: Google DeepMind / Stanford University / Hoku Labs)

Historically, these policies were based on rules developed by engineers who hard-coded them into the robot’s software—thousands of lines of code that would determine each millimeter of a robot’s movements in hundreds of tasks. What’s been happening for the last few years—and what is largely responsible for the optimism about generalist robots—is that robot policies are being handed over to advanced AI systems instead of being coded into the robot’s software. This first happened with VLMs, or vision-language models. These are similar to large language models, but they’re trained on images as well as words. Show a VLM a picture of a coffee spill and ask it to find a tool to clean up the mess, and it can identify a nearby cloth. This sort of immediate contextual understanding didn’t exist a couple of years ago when robot policies were hard-coded. 

Next came vision-language-action models, which enable robots to assess their environment and take action within it. The models do this by adding yet another component: motion commands. VLAs are trained on a series of images or videos related to performing a given task along with associated data about how a robot arm moves to perform it. That movement data is typically collected through teleoperation, in which a human uses remote controls to lead a robot through an action. This sort of training allows the AI to learn how to command the robot to move and operate during a given task. Place a VLA-powered robot in front of a desk and tell it to “close a laptop” or “wrap up the headphone wire,” and it will survey the scene, identify the relevant object, plan a way to execute the request, and then swing its arms into action—at least if it has seen this task accomplished before. 

The Gemini Robotics model is a VLA, trained on many hours of human demonstrations depicting a vast array of different actions. As a result, it can perform relatively complex tasks like picking up snow peas with kitchen tongs, doing origami, or putting together a simple lunch. It’s impressive, but there’s a glaring limitation: For now, if a robot controlled by a VLA is asked to perform a task that falls outside its training set, it’s highly likely to fail. 

“Thinking about the space of all tasks, a real generalist policy would be able to do everything along that spectrum,” says Edward Johns, a robotics professor at Imperial College London. Today, though, a Gemini Robotics model can do only “a few things here and a few things there.”

The search for true generality

So how do we get robots to be able to do more things? The usual answer is probably not surprising: Train them on more data.  

More data, the thinking goes, equals more examples, and more examples equals more generality. Google DeepMind, for instance, wants to pull together “as much data as possible,” says Pannag Sanketi, a former tech lead in robotics at the company who’s currently working on his own AI robotics project. But where to get it? Large language models had the benefit of oceans of existing text for training. There is no corresponding pool of high-quality physical demonstrations on which to train robots. Researchers have a few ways to make up for this, but all have flaws. One is to employ large numbers of people to create and collect teleoperation data (costly and time-consuming). Another is to train VLAs on videos of people performing activities (the resulting data quality is poor). Yet another is to deploy robots in the real world and use data collected from those experiences to further refine AI models (robots aren’t safe or reliable outside labs). Sanketi thinks a “multi-prong” approach that uses data collected from all these sources is the most likely path forward.

But the belief that training data alone is the answer is far from universal. Agility’s Hurst describes it as “a fundamentally flawed premise.” 

The issue is that tasks in the real world quickly explode in complexity. If you’re trying to, say, make coffee, there are myriad variables: No two kitchens are identical; coffee machines work in different ways; different cups require different grips; coffee grounds, hot water, and milk all need to be handled differently. Even this simple task requires understanding an ever-changing menu of possibilities. Achieving generality through VLAs, Hurst argues, would require “complete data coverage of all of the things that [a robot] could ever do.” Or, in other words, an almost infinite pool of training data. 

LeCun is dismissive of the whole approach. “The [AI] approaches that have been successful for language do not work for high-dimensional, continuous, noisy data”—the kind of data that is commonplace in robotics, he said in Davos. “You have to use something else.”

The leading contender for “something else” is the so-called world model—a form of AI trained less on text than on a combination of video, three-dimensional scans, and sensor data and built to predict the outcomes of actions in the real world. The aim is to build models that possess an internal representation of reality precise enough to capture how the physical world actually operates—how objects move, collide, fall, and deform. If roboticists could train machines in simulations faithful enough to real-world physics, development would become faster, cheaper, and safer, reducing the need for real-world testing. Even more transformative, robots equipped with world models could reason about their surroundings rather than merely reacting to them, helping them anticipate the consequences of an action before taking it.

Google DeepMind’s latest AI models for robots are increasingly dextrous, if rather slow and erratic

Companies like Nvidia and Google are working on the technology, and investor cash is pouring into high-profile startups. World Labs, cofounded by the Stanford AI researcher Fei-Fei Li, raised $1 billion in funding in February and was acquired by AMD at the end of September for $8.2 billion. AMI Labs, cofounded by LeCun (formerly Meta’s chief AI scientist), also raised $1 billion in March. Yet by their own admission, it is still early days. Late last year Li described the field as “nascent,” adding that “foundational approaches are still being established.” In a June Substack she described daunting challenges. For now, world models are a promising area of research rather than an immediate route to general-purpose robotics, but we are beginning to see glimmers of what they could achieve.

One such glimpse came with a small but potentially significant leap forward that took place in a San Francisco robotics lab last April. 

A breakthrough?

In the heart of San Francisco’s Mission District, the startup Physical Intelligence—or PI (as in π), as it likes to be known—is focused on developing a universal brain that could, theoretically, turn any robot into a generalist. Using an everything-including-the-kitchen sink approach to training AI models for robots, the company recently observed a hint of what a robotic brain equipped with a world model could be capable of. 

In 2024, PI published details of its first generalist robotics system, called π0, a VLA it claimed was the “most capable and dexterous generalist robot policy to date.” The model was initially trained on a 10,000-hour proprietary collection of human demonstrations gathered through teleoperation as well as several open-source robot datasets. A version released in spring 2025, π0.5, was trained on a wider variety of datasets, including labeled images from the web, lending it more versatility. A fall 2025 update, π0.6, added reinforcement learning to the model.  

Each update yielded important improvements to the model’s performance, increasing its menu of abilities from slowly folding laundry to putting things away in new environments to completing tasks like folding boxes with a higher success rate. Then, in April 2026, π0.7 seemed to catapult PI into new territory. This version makes use of a less powerful world model that generates images of steps necessary to perform a task. As the robot undertakes the job, this “lightweight” model feeds it snapshots of what to do next. 

How do you teach a robot to use a knife? At the startup Physical Intelligence, it begins with designing the right AI architecture, which includes components dedicated to language, vision and motion. This will help it relate commands — “hey robot, chop my vegetables!” — to appropriate actions.
WINNI WINTERMEYER
Data to train the AI can come from many sources, but one of the most important is human demonstrations. An employee at the startup controls a robot arm through teleoperation, exposing the AI to the task of slicing a zucchini.
WINNI WINTERMEYER