Robots Are Getting Smarter, but Your Household Helper Is Still Years Away

robots are getting smarter but your household helper is still years away Tesla's Optimus humanoid could reach public buyers by the end of 2027, according to Elon Musk. At the same time, the most sophisticated robot brain at Google DeepMind still cannot look around a kitchen and gather the ingredients for a mushroom risotto into a basket.

Tesla's Optimus humanoid could reach public buyers by the end of 2027, according to Elon Musk. At the same time, the most sophisticated robot brain at Google DeepMind still cannot look around a kitchen and gather the ingredients for a mushroom risotto into a basket.

Taken together, those two facts sum up where robotics stands today. The promises have run ahead of the machines, and the machines are still far from your kitchen.

Chances are you’ve already come across Optimus. The robot has a white body with a black head and torso, and it appears in videos dancing, handing out popcorn, dropping trash into a bin, vacuuming and tapping a microwave button. Other footage shows it toppling backward while distributing water bottles and having a hard time ironing a shirt.

The promises are huge

In Musk’s view, Optimus will be “not just Tesla’s biggest product ever, but probably the biggest product ever.” His plan is to put it to work in factories first and bring it into homes afterward. Speaking to shareholders in July, he said it eventually “will have human and then superhuman dexterity,” and he has claimed the robots could take over almost all human labor, from carrying sheet metal to folding laundry, at a price as low as $20,000 apiece. His end-of-2027 forecast came in January at the World Economic Forum’s annual meeting in Davos, Switzerland.

Musk isn’t alone. Marc Andreessen, cofounder and general partner of venture capital firm Andreessen Horowitz, has said robotics could turn into the “biggest industry in the history of the planet.” Nvidia CEO Jensen Huang said in January that humanoid robots would reach human-level ability this year. Morgan Stanley projected that robots that “resemble and act like humans” will probably number close to 1 billion by 2050, producing a market valued at more than $5 trillion.

Tesla Optimus — AI breakthroughs in robotics won't change your daily life any time soon
Robots Are Getting Smarter, but Your Household Helper Is Still Years Away 31

The reasoning runs like this: the AI behind OpenAI’s ChatGPT and Anthropic’s Claude learned to mimic human language, so a new wave of robots ought to be able to learn to mimic human movement in the same fashion.

Plenty of robotics researchers aren’t convinced. Speaking at a separate Davos event in January, Yann LeCun, frequently described as one of the godfathers of AI, said none of the companies making humanoid robots knows how to make them intelligent enough to be useful.

According to Jonathan Hurst, cofounder and chief robot officer at Agility Robotics and a robotics professor at Oregon State University, people tend to conflate two different things. “It’s very easy to make a robot that looks like a person,” Hurst said. “It is dramatically more difficult to make a machine that moves or behaves dynamically or physically like a person.”

The difference is important. A humanoid robot simply needs to resemble a human. A generalist robot needs to learn and perform many different tasks. The hype lumps the two together, and in labs across the country, debates about timelines are overshadowing progress that is gradual but genuine.

A pair of arms and a lunchbox

To watch one of today’s most capable robot brains in action, forget the humanoids. Turn instead to ALOHA 2, which stands for “A Low-cost Open-source Hardware System for Bimanual Teleoperation.”

The setup is minimal: two arms, a few grippers and a couple of cameras mounted on a bench top. It represents one side of a long-running debate among roboticists. Fans of humanlike robots argue that the human form will let machines slot into a world that was built for people. Skeptics argue it isn’t worth the effort. ALOHA 2 is the skeptics’ response, and Google DeepMind researchers rely on it in their labs to test Gemini Robotics, their most advanced AI system for robots.

Tesla Optimus — AI breakthroughs in robotics won't change your daily life any time soon
Robots Are Getting Smarter, but Your Household Helper Is Still Years Away 32

Under the control of Gemini Robotics, ALOHA 2 begins to behave more like a generalist, able to tackle any number of tasks it has encountered examples of during training. In a video showing a lunch-packing exercise, it uses two pincer grippers to ease a slice of white bread into a Ziploc bag and seal it. It then places a bunch of grapes into a Tupperware container, fastens the lid, carefully transfers everything into a lunchbox and zips it closed.

It’s hardly an appetizing meal. Even so, a robot assembling it unaided marks a clear advance over what could be done just three years ago.

Much of that advance stems from AI’s impact on robot policies. A policy is the component that determines how a general-purpose robot interprets its environment, plans its motions and then executes the task properly.

From hand-written code to models that learn by watching

Policies used to be coded manually by engineers, which meant thousands of lines specifying every millimeter of a robot’s motion across hundreds of tasks. In recent years, that job has moved to advanced AI systems, and the change accounts for most of today’s optimism about generalist robots.

The first step was vision-language models, or VLMs. These function like large language models but learn from pictures in addition to text. Give a VLM an image of spilled coffee, ask it to locate something to wipe up the mess, and it can identify a cloth nearby. Back when policies were hard-coded a couple of years ago, robots had no such understanding of context.

Tesla Optimus — AI breakthroughs in robotics won't change your daily life any time soon
Robots Are Getting Smarter, but Your Household Helper Is Still Years Away 33

Vision-language-action models, or VLAs, layer motion commands on top. A VLA is trained on images or video of a task paired with data describing how a robot arm moves to complete it. That motion data typically comes from teleoperation, in which a person guides a robot through the movement using remote controls. Place a VLA-driven robot in front of a desk, instruct it to “close a laptop” or “wrap up the headphone wire,” and it will survey the scene, locate the correct object, plan the motion and set its arms moving. This works provided it has seen the task performed before.

Gemini Robotics is a VLA trained on many hours of human demonstrations spanning a broad array of actions. That training is what lets it grab snow peas with kitchen tongs, fold origami or put together a basic lunch.

The catch is significant. Ask a robot running a VLA to attempt something beyond its training data, and failure is very likely.

“Thinking about the space of all tasks, a real generalist policy would be able to do everything along that spectrum,” said Edward Johns, a robotics professor at Imperial College London. Right now, Johns said, a Gemini Robotics model can handle only “a few things here and a few things there.”

An unsolved data problem

The industry’s usual response is no surprise: feed the models more data, meaning more examples, which in theory yields more generality. Pannag Sanketi, formerly a tech lead in robotics at Google DeepMind and now working on his own AI robotics project, said the company aims to collect “as much data as possible.”

Getting that data is the tricky part. Large language models could draw on vast stores of existing text. There’s no equivalent pool of high-quality physical demonstrations.

Each alternative comes with trade-offs. Hiring many people to generate teleoperation data is costly and slow. Training VLAs on footage of people performing tasks yields low-quality data. Deploying robots in the real world to gather experience collides with the fact that robots aren’t safe or dependable outside the lab. Sanketi said a “multi-prong” strategy drawing on all of these sources is the likeliest route forward.

Hurst doesn’t believe more data solves the problem. He described the notion as “a fundamentally flawed premise.”

In his view, real-world tasks become complex very quickly. Consider brewing coffee. No two kitchens are alike, coffee makers operate differently, cups require different grips, and grounds, hot water and milk all need their own handling. Reaching generality with VLAs, Hurst said, would demand “complete data coverage of all of the things that [a robot] could ever do.” In practice, that means a nearly endless supply of training data.

LeCun put it more bluntly at Davos. “The [AI] approaches that have been successful for language do not work for high-dimensional, continuous, noisy data,” he said, referring to the sort of data robots constantly encounter. “You have to use something else.”

Investors are betting on world models

The leading candidate for that “something else” is the world model. This type of AI is trained less on text and more on video, 3D scans and sensor data, and it is designed to forecast the outcome when something acts in the physical world. The aim is an internal representation of reality precise enough to capture how objects move, collide, fall and change shape.

That would pay off in two ways. If simulations reproduced real physics faithfully enough, robots could be trained inside them, making development quicker, cheaper and safer while reducing real-world testing. A robot equipped with a world model could also reason about its environment rather than merely react to it, anticipating the result of an action before carrying it out.

Nvidia and Google are both developing the technology, and investors are funneling money into prominent startups. World Labs, cofounded by Stanford AI researcher Fei-Fei Li, raised $1 billion in February and was bought by AMD at the end of September for $8.2 billion. AMI Labs, cofounded by LeCun, Meta’s former chief AI scientist, likewise raised $1 billion in March.

Even the field’s leading figures say it’s early days. Late last year, Li called the field “nascent” and said “foundational approaches are still being established.” In a June Substack post, she outlined major challenges. For now, world models are a promising line of research rather than a shortcut to general-purpose robots, although some early results are beginning to hint at their potential.

A sweet potato, almost air-fried

The most intriguing of those results emerged in April from a lab in San Francisco’s Mission District. Physical Intelligence, a startup known as PI (as in π), is trying to build a universal brain that could, in principle, transform any robot into a generalist. Its strategy is to train on every source it can obtain.

PI released details in 2024 of π0, its first generalist robotics system, a VLA the company described as the “most capable and dexterous generalist robot policy to date.” π0 was initially trained on a proprietary collection of 10,000 hours of teleoperated human demonstrations along with several open-source robot datasets. π0.5 followed in spring 2025, trained on a wider blend of data that included labeled web images, making it more versatile. An update in fall 2025, π0.6, introduced reinforcement learning.

Every iteration improved. The model progressed from slowly folding laundry, to tidying items away in unfamiliar settings, to folding boxes with a higher success rate.

π0.7 arrived in April 2026. It relies on a less powerful world model that produces images of the steps a task requires. As the robot works, this “lightweight” model supplies it with snapshots showing what to do next.

According to PI, the model displays the first signs of compositional generalization, which is when an AI system carries out a skill it has never encountered by recombining skills from its training data. In one trial, PI instructed the model to “load a sweet potato into the air fryer,” a task it had never seen. The demo video shows the robot fumbling somewhat and making a few false starts before ultimately putting in a reasonable effort, though it doesn’t quite complete the job.

Sergey Levine, a University of California, Berkeley professor and PI cofounder, is excited by the outcome. “It’s actually the first time that we’ve convincingly seen that kind of compositional generalization, where we can basically ask the model to do tasks that we did not specifically collect data for and train it to do, and it’ll actually make a passable attempt,” Levine said.

What happened next is arguably the most telling part. The team combed through the training data to understand how the model pulled it off. They uncovered fragments of relevant labeled teleoperation data, among them two instances of a human operator using the robot to slide an air fryer basket into the fryer. Those two snippets may have been sufficient to carry π0.7 most of the way toward cooking a sweet potato.

Just how impressive π0.7’s generalization is remains uncertain. Yet given that the model only glimpsed air fryers briefly, the outcome gives a rough sense of how far cutting-edge research can push robots right now.

Who’s really pulling the strings

There’s a noticeable gulf between these modest lab victories (“Look! It put a sweet potato into an air fryer!”) and the lifelike nimbleness of viral clips showing robots dancing onstage and courteously serving drinks. Those clips frequently omit a crucial detail: in many cases, a person is operating the robot or has meticulously scripted its actions.

The robot that shared the stage with Huang in March 2025, apparently obeying his commands and following him around, was being remotely controlled by what its creators called “a puppeteer behind the scenes.”

Fully autonomous motion planning, in which a robot independently figures out where to go, remains a largely unsolved challenge, particularly in novel and chaotic settings such as a construction site or an unfamiliar house. A larger and often connected challenge is getting robots to tackle bigger, less defined jobs composed of multiple tasks, where the robot must determine how to handle each one and in what sequence. A robot capable of placing a plate in the microwave is far removed from one that hears “Make dinner,” inspects the refrigerator, chops ingredients and switches on the stove. The risotto test described at the outset is Google DeepMind’s effort at exactly that kind of job, and it hasn’t succeeded yet.

Reliability poses yet another hurdle. To be useful, a robot needs to succeed virtually every time. Marc Raibert, founder of Boston Dynamics, said VLA researchers are cheering the wrong milestones. “People are very excited when their result goes from 50% success to 70% success,” Raibert said. “But 70% success is like it doesn’t work, right?”

What works in the real world is mundane

The handful of humanoids undergoing real-world trials perform narrowly defined tasks in tightly controlled settings. Not one of them is a generalist.

Agility said it has hundreds of robots in trials at facilities run by GXO Logistics, Amazon and Schaeffler. Currently, Hurst said, they perform simple jobs such as moving bins and totes. He added that it took years to make the robots safe enough for logistics firms to even consider deploying them.

Musk’s history on this front deserves scrutiny. In May 2025, he asserted that “thousands” of Optimus robots would be at work in Tesla factories by the close of that year. In January of this year, he said the company had just “some of the Tesla Optimus robots doing simple tasks in the factory.”

The $20,000 home robot still relies on a person

If factories are tough, homes are tougher. The 1X Neo home robot is available for preorder now and is slated to ship at some point later this year for $20,000. It pledges to take on “the boring and mundane tasks around the house,” such as putting dishes away, answering the door and straightening up the living room, “so you can focus on what matters to you.”

The goal is for the five-foot-six-inch robot to eventually handle all of this on its own. Currently, though, most tasks require a remote human operator, meaning you’d need to let a stranger view your home through the robot’s cameras.

Asked how long it will be before fully autonomous robots are ready for household chores, Hurst offered an estimate. “If I had to pick a number, I’d say it’s 10 years before robots are … actually doing useful things in people’s homes,” Hurst said.

By the time that happens, the robots could very well be Chinese. Chinese companies accounted for nearly 90% of the approximately 15,000 humanoid robots shipped in 2025, according to market intelligence firm Omdia and Chinese robotics company Unitree. Unitree shipped more humanoids than any rival last year, and one of its models sells for under $6,000, beating Musk’s most optimistic Optimus price by more than $14,000, although Unitree anticipates its machines will enter industrial work first. The AP recently reported that the buyers are primarily corporate and academic labs along with state-owned enterprises.

A familiar demo

The pursuit of humanoid robots goes back a long way. Leonardo da Vinci drew a mechanical knight driven by cables and pulleys in 1495. At the 1939 New York World’s Fair, Westinghouse’s Elektro, a seven-foot-tall box on legs, smoked a cigarette. In 1973, Japan’s Waseda University built WABOT-1, the first full-scale programmable humanoid.

Honda’s ASIMO, introduced in 2000, was likely the first that showed any real competence. Up to a point, it could climb stairs, recognize faces and navigate spaces independently. Honda discontinued it in 2018 because it never progressed far enough beyond its demos to be practical.

Each of those machines dazzled audiences in its time, and none could cope with the real world. The air fryer experiment tells the same tale. For all the AI powering them, today’s robots still stall at the step that would carry them from a scripted demo into a cluttered kitchen they’ve never encountered.