- Published on
Figure’s Helix 2.5 Takes on Chores in 30 Unseen Homes
One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

Supported by RoboStrategySupport for Humanoids Daily comes from RoboStrategy
SponsoredInvesting involves risk, including possible loss of principal. Read the prospectus before investing.
- Figure says Helix 2.5 completed household chores in 30 previously unseen homes, without training or adaptation in those homes.
- Its Index-pretrained policies completed 237 of 420 trials, about 56%, across bed making, towel folding and tidying toys.
- “Zero-shot” applies to the homes and manipulated objects. The three behaviors were learned through task-specific training elsewhere.
- The company-run evaluation supports transfer to unfamiliar settings, while leaving a substantial reliability gap before dependable household service.

Figure has introduced Helix 2.5, reporting that its humanoids can make beds, fold towels and tidy living rooms in homes they have never encountered during training. In a September 17 announcement, the company describes testing across 30 Bay Area homes, with its strongest policies completing about 56% of full-task trials.
The significance is the change of setting. A household robot becomes much more useful if learning a chore once allows it to perform that chore in another home, without collecting a fresh training dataset for every customer’s furniture and belongings.
CEO Brett Adcock said Figure rented the homes for the experiment. He also posted nearly four hours of extended footage showing Figure 03 working in different homes, giving viewers more material to inspect beyond the announcement’s shorter demonstrations. The footage is not, by itself, a complete record of the 420 evaluated trials.
In the accompanying video, Adcock calls Helix 2.5 “the most important project we’ve ever taken on at Figure.” He and Director of AI Corey Lynch say the robot recorded successes in every one of the 30 homes. That describes the reach of the results, rather than a perfect completion rate: the pooled trial success rate was still 56%.
The weekly humanoid robotics briefing
One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.
Read recent issuesNew homes, previously learned chores
Helix 2.5 starts from a foundation model pretrained on Index, Figure’s dataset of human behavior. That base was then adapted to three separate behaviors. Each task used one fixed checkpoint across the 30 homes; this is not a claim that an untrained robot arrived and learned arbitrary chores on the spot.
Figure says no data was collected in the evaluation homes and none of the evaluation toys, towels or bedding appeared in task-specification data. Robots used the homes’ existing beds, couches and folding surfaces. Evaluation performance was not used to select checkpoints, according to the company.
The tasks required movement as well as manipulation: searching different surfaces for toys, positioning around furniture and walking around beds while handling bedding.
What the robots completed
Figure’s results chart reports the following for the Index-pretrained policies:
| Task | Successful trials | Success rate |
|---|---|---|
| Bed making | 94/140 | 67% |
| Towel folding | 87/140 | 62% |
| Tidying toys | 56/140 | 40% |
| All three combined | 237/420 | 56% |
Success required finishing the entire task, rather than receiving partial credit for individual objects. Tidying meant collecting all 13–15 toys; towel trials required folding and placing all four towels in a basket. Safety interventions counted as failed trials, and the rubric imposed timeouts.
To test Index’s contribution, Figure compared policies with identical task-specific data, architecture and downstream training, varying whether they received Index pretraining. Its chart lists 35 successful trials out of 420 for the baseline, about 8.3%. The announcement’s prose gives that baseline as 9%, a small discrepancy in the published figures.
Either way, the reported improvement is substantial. It is also a company-run comparison, not independent validation, and 56% completion still leaves roughly 44% of trials unfinished under the evaluation criteria.
Sunday’s Tony Zhao questions the usefulness claim
Update, September 17: Added Zhao’s response and context from Sunday’s ACT-2 evaluation.
Sunday Robotics’ Tony Zhao challenged Figure’s framing, citing the 237 successful trials out of 420. He described the result as “failing half the time” and argued: “Doing useful work = generalization + reliability.” The exact full-task failure rate is about 44%; a failed trial does not necessarily mean the robot made no useful progress.
The criticism follows Sunday’s own ACT-2 announcement. Sunday reported 778 successful folds across 785 autonomous attempts, or 99.1%, spanning nine garment types in unseen environments, with no per-home adaptation. Its technical write-up defines success as autonomously folding and stacking an individual garment and grades fold quality separately.
The percentages are not directly comparable. Sunday counts individual garment attempts; Figure’s pooled result covers complete bed-making, tidying and towel-folding trials. Figure’s towel-only result is 87/140, about 62%, but each trial requires all four towels to be folded and placed in a basket. Different objects, task lengths and grading rules prevent those figures from establishing a head-to-head ranking.
Zhao’s objection concerns the threshold for useful household work. Figure’s experiment addresses whether learned behaviors transfer to unfamiliar homes, and the company itself says general humanoid robotics remains unsolved. Both releases are company-reported evaluations; together, they underline why claims about generalization need clearly defined measures of reliability.
The payoff Figure wants from Index
The release gives a concrete test of the strategy behind Figure’s crowdsourced Index platform: use broad human experience to reduce how much must be taught directly on robots.
Lynch says in the video that over 90,000 people now contribute to Index each week, with 35 minutes of human experience uploaded every second. The latter rate also appears in Figure’s written announcement. These are company-reported collection figures; contributor growth alone does not establish the quality or usefulness of the resulting training data.
Figure also reports predictable improvements in robot-action prediction across four training runs spanning an eightfold increase in Index data. That measures prediction loss at a fixed model size, rather than proving that real-world success will keep rising at the same rate. The distinction matters as Figure expands training through its multi-billion-dollar Nscale compute agreement.
The next commercial question is how consistently those learned behaviors survive everyday use: interrupted chores, occupied rooms and repeated work over days. Helix 2.5 makes unfamiliar homes a more credible test environment for humanoids. Dependable service still requires substantially fewer failed tasks.
Share this article
Read next
- Published on
- Reading time
- 4 min read
Figure AI Reports Rapid Growth for Index, Surpassing 69,000 Weekly Active Users
- Published on
- Reading time
- 4 min read
Figure AI Unveils Index: Crowdsourcing Real-World Human Video to Train Helix
- Published on
- Reading time
- 5 min read
BMW Reflections: Why the Spartanburg Pilot Succeeded and Why Figure 02 Had to Die
- Published on
- Reading time
- 4 min read
The End of C++: Brett Adcock on Helix 02 and Figure’s Path to "Room-Scale" Autonomy
- Published on
- Reading time
- 3 min read
The Dishwasher Wars: Figure Fires Back at Sunday Robotics with Glassware Demo
- Published on
- Reading time
- 5 min read
From Pixels to Torque: Figure Unveils Helix 02 and the Era of Whole-Body Autonomy
The weekly humanoid robotics briefing
One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.
Read recent issues
















