Helix 2.5: one model, three chores, and 30 homes the robot had never entered
Updated on September 18, 2026
Helix 2.5 is the AI model that drives Figure's humanoid robot through whole-body chores, from walking across a room to folding laundry. Figure revealed it in September 2026 with an unusual field test, 30 rented Bay Area homes where the robot tidied toys, folded towels and made beds with zero on-site training. It completed 56% of 420 attempts, against 9% for the same system without pretraining. Nothing is for sale at this point, neither the model nor the robot.
- Zero-shot results measured in 30 real homes
- One network for walking, seeing and grasping
- Recovers and retries after its own failed grabs
- Half the task data required by Helix 02
- 56% success, still far from daily reliability
- No weights, API or product to buy yet
- Research-stage system, watch-only for now
Three chores, 30 living rooms, no partial credit
Figure rented 30 homes and dropped its robot in without collecting a single frame of data on site, which the company describes as a first at this scale for a humanoid. The tasks were tidying a living room scattered with 13 to 15 toys, folding towels and making a bed, using each home's own furniture and linens.
Grading left no wiggle room. A chore counted only if finished end to end, which yields 237 completed trials out of 420, with at least one success in every single home. Picture the run, the robot walks until a toy enters its view, shifts its stance, grabs with both hands, then circles the bed to redo a sloppy fold.
Index does the heavy lifting for Helix 2.5
Pretraining on Index, Figure's dataset of human-behavior video, accounts for most of the result. With task data, architecture and training settings held identical, zero-shot success jumps from 9% to 56% once Index enters the picture. There is a second break with the past, Helix 02 started from an existing vision-language model, while Helix 2.5 was pretrained from scratch on this human data alone.
The scale is dizzying. Index now collects roughly 35 minutes of fresh human experience every second, and Figure has committed $3.5 billion of compute to training Helix.
| Metric | Verified figure |
|---|---|
| Homes tested with no adaptation | 30 |
| Completed trials | 237 out of 420, or 56% |
| Same policy without Index | 9% success |
| Task data versus Helix 02 | Half as much |
| Fresh Index data | 35 minutes per second |
Helix 2.5 is a preview, not a product on a shelf
No model weights, API or pricing have been published for Helix 2.5, and the Figure 03 robot carrying it cannot be bought by the public either. Home pilot deployments are planned, with a consumer price targeted below $20,000, though nothing official has been confirmed yet.
For now, everything happens from the couch. Figure CEO Brett Adcock posted close to four hours of raw footage from those houses (failed grabs included, which makes a nice change from overly polished demos). Among rivals, 1X's NEO leans on human teleoperation for tricky tasks and Tesla keeps refining Optimus, two different roads from the full autonomy claimed here. Like most home robots and AI devices of this generation, this one is a story to follow rather than a box to order.
Frequently asked questions
What can Helix 2.5 do?
Three household chores so far, tidying a living room, folding towels and making a bed, all performed by the Figure 03 robot in homes absent from its training data. The model handles walking, perception and two-handed coordination at once, and repositions itself when a grasp goes wrong.
Can you buy the Figure 03 robot running Helix 2.5?
No, not yet. Figure sells nothing to the public and has opened no pre-orders. A consumer price below $20,000 has been floated without confirmation, and pilot deployments in selected homes are planned before any wide release. Registering interest on figure.ai is the only step available today.
Is Helix 2.5 really zero-shot?
Yes, in the sense Figure defines. No data was gathered in the 30 evaluation homes and none of the manipulated objects appeared in training data. The three behaviors themselves were still taught elsewhere through demonstrations. The term covers the places and objects, not the initial learning of the skills.
What is the difference between Helix 2.5 and Helix 02?
The starting point changes everything. Helix 02 built on a pretrained vision-language model, while Helix 2.5 was pretrained from random initialization purely on Index, Figure's collection of human-behavior video. It matches a Helix 02 policy with half the task-specific data, and generalizes across 30 homes it never saw.
Verdict: A robot that makes a stranger's bed roughly half the time may sound modest, yet it is the strongest sign so far that a chore learned once can travel to new places. Keep Helix 2.5 on your radar if you follow home robotics, embodied AI research or the slow march of humanoids toward real houses.
