Figure AI released Helix 2.5 on September 17, 2026 and reported that its humanoid did household chores (tidying toys, folding towels and making beds) in 30 Bay Area homes it had never seen, with no data collected in those homes and no weights adapted after deployment. The company reports a 56% success rate for the model pretrained on Index, its dataset of human behavior, against 9% for a model trained without that pretraining.

Why it matters: home robots such as 1X's NEO have to work in houses their makers have never mapped. Figure has now put a multi-home, zero-shot success rate on that problem, and the number shows how far the task still is from reliable.

What did Helix 2.5 do in the 30 homes?

Figure evaluated three tasks with strict pass criteria. Living-room tidying counted as a success only if all 13 to 15 toys scattered in the scene were picked and placed in a basket. Towel folding required all towels to be folded and placed in the basket. Bed making required both pillows and the comforter corners to be placed at the top of the bed. Figure says none of the evaluation objects appeared in the training data. The company did not name the robot model in its post, and did not state the exact number of trials.

The headline comparison is 56% zero-shot success for the Index-pretrained model versus 9% for the baseline, a roughly sixfold difference attributed to pretraining alone. All results are company-reported.

Key Facts

  • Released September 17, 2026; evaluated in 30 unseen Bay Area homes
  • Tasks: living-room tidying (13–15 toys), towel folding, bed making
  • 56% zero-shot success with Index pretraining vs 9% without (company-reported)
  • Half the task-specific data of a representative Helix 02 behavior
  • Index generates about 35 minutes of human experience per second; no evaluation task exceeds 1.90% of it

What is Index, and how is Helix 2.5 different from Helix 02?

Figure describes Index as its "global-scale dataset of human behavior" and says it is now generating roughly 35 minutes of new human experience every second. It has not published how the data is collected. Helix 2.5 was pretrained from random initialization entirely on Index, whereas Helix 02 started from a pretrained vision-language model. Figure says Helix 2.5 used half as much task-specific data as a representative Helix 02 behavior, and that no single evaluation task makes up more than 1.90% of the Index pretraining dataset.

What is Figure's scaling-law claim?

Figure reports that loss "fell predictably with each doubling of Index," with forecasting error of 0.54% of the variation across the full 8x data range. Using only its smaller runs, the company says it predicted its largest run's test loss "to four decimal places before training began." The AI Insider reports that Figure called this the first human-to-robot transfer scaling law measured on a humanoid.

A predictable curve is what justifies large compute purchases in advance. Figure and Nscale announced an initial $3.5 billion compute commitment on September 3, covered in our report on the deal. The 56% figure also means that, on Figure's own strict criteria, 44% of zero-shot attempts did not finish the task.

Frequently Asked

What is Figure Helix 2.5?

Helix 2.5 is Figure AI's robot foundation model, released on September 17, 2026. It is pretrained from scratch on Index, Figure's dataset of human behavior, and Figure reports it did household chores zero-shot in 30 unseen Bay Area homes.

How successful was Helix 2.5 in unseen homes?

Figure reports 56% zero-shot success across tidying, towel folding and bed making for the Index-pretrained model, versus 9% for a model without that pretraining. A trial counted only if the whole task was finished, such as all 13 to 15 toys placed in the basket.

What does zero-shot mean for Helix 2.5?

Figure says no data was collected in the 30 evaluation homes, none of the evaluation objects appeared in the training data, and no model weights were adapted after deployment.

How is Helix 2.5 different from Helix 02?

Helix 02 started from a pretrained vision-language model, while Helix 2.5 was pretrained from random initialization entirely on Index. Figure says Helix 2.5 used half as much task-specific data as a representative Helix 02 behavior.