Figure's Helix 2.5 Walked Into 30 Homes It Had Never Seen
Index pretraining → 56% zero-shot success vs 9% from scratch. Tidying, towels, beds — no fine-tuning in those houses. Homes aren't the demo. Transfer is.
Most humanoid demos are a costume change for the same room. Figure’s September 17 write-up asks a meaner question: can a robot walk into a house it has never mapped and still tidy the living room, fold towels, and make the bed — whole body, zero shot?
So basically: Helix 2.5 isn’t interesting because it can do chores. It’s interesting because Index pretraining moved zero-shot success from 9% (train from scratch) to 56% — same task data, same architecture, same eval — across 30 unseen Bay Area homes.
What they actually tested
Figure pretrained Helix 2.5 on Index, its global-scale human-behavior dataset, then adapted one foundation model into three long-horizon behaviors: tidying living rooms, folding towels, and making beds. Evaluation homes and manipulated objects were held out — no data collection, fine-tuning, or adaptation in those environments.
Success meant finishing the whole task: every toy in the basket, every towel folded into the basket, bed made to pre-fixed criteria. No partial credit. One checkpoint per task across all 30 homes.
Holding task-specific data fixed, Index pretraining alone drove that 9% → 56% jump. Figure also says Helix 2.5 used about half as much task-specific data as a representative Helix 02 behavior, then generalized across thirty homes instead of one trained room.
Why “homes” are the hard part
Tabletop arms get a fixed workspace. Wheeled bases get open floor. Homes give you tight furniture, weird layouts, and objects that weren’t in the fine-tune set. Perception, locomotion, and manipulation stop being separable — the robot has to walk to see, then shift stance to reach.
Figure leans hard on a scaling story: Index is generating roughly 35 minutes of new human experience per second, and the company says it has committed $3.5B of compute to training Helix. They also claim a human-to-humanoid transfer scaling law — doubling Index data predictably improved downstream robot-action prediction.
What to watch
56% zero-shot is not “solved homes.” It’s evidence that pretraining on human experience transfers better than rebuilding every kitchen from scratch. Watch for:
- Independent evals that aren’t Figure’s own rubric
- Failure modes in cluttered / multi-pet / multi-floor homes
- Whether the $3.5B compute bet shows up as public capability, not just blog charts
The grounded take
The demo that matters isn’t a robot making a bed. It’s a policy that doesn’t need a week of data collection every time the couch moves three feet. Helix 2.5 is Figure arguing that whole-body intelligence can be learned from people first, then specified once — and that the next doublings of Index should move the needle in a measurable way.
Treat the press release as a progress report on transfer, not a shipping date for a housekeeper. The filter is still the same: boring reliability in someone else’s living room on a Tuesday.
So basically — pass it on.