Helix 2.5, Figure’s new AI model for humanoid robots, has landed on AIxploria. The directory published its card at 04:59 UTC on 18 September 2026, about sixteen hours after Figure announced the model, with a 4.5 rating from two votes, a “Verified Tool” label, a “Paid” tag and a badge reading “#4 in Robots and Devices”. By early afternoon it had 7,398 views and 63 upvotes.
The model itself is a genuine research result. Figure took one humanoid into 30 Bay Area homes it had never seen, collected no data in any of them, and had it tidy toys, fold towels and make beds. With its new pretraining the robot succeeded 56 per cent of the time; the same system trained from scratch managed 9 per cent.
We checked the card the way this series checks every listing: claim by claim against Figure’s own announcement and related pages, with the badge walked against the category it links to and the ranked lists searched for the new entry. The numbers mostly hold. The labels are another matter: a model nobody can buy is tagged “Paid”, and a piece of software now ranks above the robot it runs on.
Table of contents
- What the Helix 2.5 Card Says
- Checking the Helix 2.5 Card Against Figure’s Announcement
- The “Paid” Tag on a Model Nobody Can Buy
- The Results Behind Helix 2.5, Task by Task
- How Strict Was the Helix 2.5 Grading?
- Index: The Dataset Doing the Work
- The Scaling Law the Card Leaves Out
- The Badge: #4 in Robots and Devices
- The Alternatives Strip and the Ranked Lists
- Helix 2.5 Against the Rest of the Field
- What Helix 2.5 Means Outside the Lab
- Helix 2.5 on AIxploria: The Verdict
- References
What the Helix 2.5 Card Says
The card is a full editorial review headed “one model, three chores, and 30 homes the robot had never entered”. It is filed in two categories, Robots and Devices and Latest AI.
The summary
The blurb calls Helix 2.5 “a humanoid robot controlled by a neural network that tidies up, folds laundry, and makes beds in homes it has never visited (without collecting data on site)”. The review adds the headline numbers: 56 per cent success over 420 attempts, against 9 per cent without pretraining, and “nothing is for sale at this point, neither the model nor the robot.”
Pros and cons
Four pros: zero-shot results measured in 30 real homes, one network for walking, seeing and grasping, recovery after failed grabs, and half the task data needed by Helix 02. Three cons: 56 per cent is “still far from daily reliability”, there are no weights, API or product, and it is a research-stage system that is “watch-only for now”.
The verdict
The review ends that a robot making a stranger’s bed roughly half the time “may sound modest, yet it is the strongest sign so far that a chore learned once can travel to new places.” That is a fair reading of Figure’s evidence, and it is a more careful verdict than most launch coverage managed.
Checking the Helix 2.5 Card Against Figure's Announcement
Figure published “Helix 2.5: Zero-Shot 30-Home Generalization” on 17 September 2026. We compared every checkable claim on the card with that post, Figure’s earlier Index and Nscale announcements, and the news reports that carried additional figures.
| Card claim | What Figure or coverage says | Verdict |
|---|---|---|
| 30 homes, no data collected on site | “30 Bay Area homes with zero data collected in any of them” | Matches |
| 56% success vs 9% without Index | “from 9% to 56%”, with all other variables held fixed | Matches |
| 237 of 420 trials completed | Not in the post’s text; reported by Startup Fortune and Crypto Briefing with a per-task split | Consistent with coverage |
| 13 to 15 toys per tidy | “All 13-15 toys scattered in the scene” | Matches |
| Half the task data of Helix 02 | “half as much task-specific data as a representative Helix 02 behavior” | Matches |
| Pretrained from scratch on Index | “pretrained from random initialization entirely on Index” | Matches |
| 35 minutes of human data per second | “roughly 35 minutes of new human experience every second” | Matches |
| $3.5 billion of compute committed | Nscale deal: “initial commitment of $3.5 billion”, intent to scale past $6 billion | Matches, context missing |
| Homes were rented | Brett Adcock: “We rented 30 homes in the Bay Area” | Matches the CEO’s post |
| Runs on the Figure 03 robot | Figure’s post does not name the robot; Crypto Briefing says Figure 03 | Not stated by Figure |
| Consumer price below $20,000 | Not on any Figure page we read | Unsourced |
| At least one success in every home | Not in the post’s text | Not found |
| “Paid” tag | The card’s own review: “Nothing is for sale” | Self-contradictory |
The score
Of thirteen claims, eight match Figure or its chief executive, one is consistent with independent coverage, two could not be found, one is unsourced and one contradicts the card’s own text. On the research itself, the Helix 2.5 card is accurate. The drift is all in the commercial details.
The price line
The card says a consumer price “below $20,000” is targeted, and hedges that “nothing official has been confirmed yet”. We found no such figure on Figure’s Helix 2.5 post, the Figure 03 launch page or the Index and Nscale announcements. With that hedge the line is honest, but a reader skimming the card will remember the number, not the caveat.
The "Paid" Tag on a Model Nobody Can Buy
Every AIxploria card carries a pricing tag, and it is drawn from the structured data behind the page. For Helix 2.5 that data says “Paid”.
What the card itself says
The same page states that “no model weights, API or pricing have been published for Helix 2.5”, that the robot “cannot be bought by the public”, and, in its FAQ, that “Figure sells nothing to the public and has opened no pre-orders”. The only available step, it says, is registering interest on Figure’s website.
Why the tag is wrong in a specific way
“Paid” implies a transaction is possible. “Free” would be equally wrong. The honest label is something like “Not available”, which AIxploria’s tag set does not appear to include. This is the same class of error the series found on the MAI-Transcribe-2 card, whose “Free” tag sat above an FAQ quoting a price.
Why it matters to readers
Directory tags drive filters. A reader browsing paid robotics tools will find Helix 2.5 beside products with checkout pages. The review text fixes the impression for anyone who reads it; the tag misleads anyone who does not.
The Results Behind Helix 2.5, Task by Task
Figure’s post gives the headline 56 per cent and a chart. The per-task numbers come from reports that read the trial counts: Startup Fortune and Crypto Briefing both give 237 successes out of 420 attempts, split across three tasks of 140 trials each.
| Task | Successes | Attempts | Success rate | What counted as success |
|---|---|---|---|---|
| Bed making | 94 | 140 | 67% | Both pillows and comforter corners at the top, comforter pulled smooth |
| Towel folding | 87 | 140 | 62% | All towels folded and placed in the basket |
| Living room tidy | 56 | 140 | 40% | All 13 to 15 toys picked and placed in the basket |
| All three | 237 | 420 | 56.4% | Full completion only; no partial credit |
The arithmetic checks out
94 plus 87 plus 56 is 237, and 237 divided by 420 is 56.4 per cent, which rounds to Figure’s 56. Three tasks of 140 trials across 30 homes is about 4.7 trials per task per home, so each home saw roughly fourteen attempts in total.
Why toys are hardest
The tidy task needs every one of 13 to 15 toys placed, with a one-minute timeout per toy. One missed toy fails the whole trial. Bed making, by contrast, involves a handful of large objects. The spread tells you more about how strictly each task was graded than about which chore is intrinsically harder.
The six-fold claim
Figure describes 56 per cent as “over 6x higher” than 9 per cent. That is correct: 56 divided by 9 is 6.2. Because only the pretraining changed between the two policies, Figure argues the gap “directly measures its contribution”.
How Strict Was the Helix 2.5 Grading?
The card mentions “no partial credit”. Figure’s appendix goes further, and it is the part of the release that makes the numbers credible.
Timeouts on every object
Each toy got a one-minute timeout, and exceeding it aborted the rollout. Each towel got three minutes. Each pillow and each side of the comforter got one minute. Graders also recorded attempts per toy and graded the quality of each fold, even when the trial failed.
Safety interventions count as failures
“If a human intervention is necessary for safety, that rollout is aborted and failed.” That closes the most common loophole in robot demos, where a hand quietly steadies the robot off camera.
Fixed checkpoints, blind evaluation
Each task used “a single fixed checkpoint across all 30 homes”, no weights were adapted to the homes, and no evaluation data was used to pick checkpoints. Evaluation objects were set aside in advance and checked by an AI model and then a human to confirm they were absent from the training data.
What “zero-shot” does and does not mean
Figure is careful here, and so is the card’s FAQ. The homes and objects were new. The behaviours were not: they were taught through fine-tuning data collected elsewhere. “Zero-shot” covers places and objects, not the skills themselves.
Index: The Dataset Doing the Work
Most of what is new in Helix 2.5 comes from Index, the human-video dataset Figure unveiled three weeks earlier. The card says Index “does the heavy lifting”, and Figure’s own numbers support that.
What Index is
Figure brought Index out of stealth on 25 August 2026 as “the most diverse robot training dataset ever built”. People record themselves doing real tasks through a phone app; Figure calls them Creators and says it had paid them $15 million by launch.
The scale, then and now
At launch Figure reported 264,000 app downloads across 108 countries, over 44,000 weekly active users, more than 16 million uploaded videos and “30 minutes of video uploads every second”. By the Nscale announcement on 3 September the rate was 35 minutes a second, and the Helix 2.5 post repeats “roughly 35 minutes”.
Thirty-five minutes a second is 35 times 86,400 seconds, or about 3.0 million minutes of video a day, roughly 5.8 years of footage every 24 hours. Figure’s own launch figure, “4.9 years of human work every day”, was set at the 30-minute rate, and 4.9 scaled by 35/30 gives 5.7, so the two statements agree.
| Date | Figure announcement | What it added |
|---|---|---|
| 9 Oct 2025 | Introducing Figure 03 | Palm cameras, fingertip tactile sensors, a home-safe design and the BotQ factory |
| 27 Jan 2026 | Introducing Helix 02 | Whole-body control; a 4-minute autonomous dishwasher task |
| 25 Aug 2026 | Introducing Index | Human-video dataset; 30 minutes uploaded per second |
| 3 Sep 2026 | Nscale partnership | Up to 100,000 Vera Rubin GPUs; $3.5bn initial compute commitment |
| 17 Sep 2026 | Helix 2.5 | Index-pretrained model; 30-home zero-shot evaluation |
| 18 Sep 2026 | AIxploria card | Published 04:59 UTC, about 16 hours after Figure’s post |
How broad the pretraining is
Figure says “no single evaluation task makes up more than 1.90% of the Index pretraining dataset.” That is an important detail the card leaves out, because it is the answer to the obvious objection that the model simply saw lots of bed making.
A different starting point from Helix 02
Helix 02 began from a pretrained vision-language model and used a whole-body controller, System 0, trained on over 1,000 hours of human motion data and sim-to-real reinforcement learning. Helix 2.5 was pretrained from random weights on Index alone. That is a bet that human video can replace internet-scale vision-language pretraining for a robot’s foundation.
The Scaling Law the Card Leaves Out
The most consequential claim in Figure’s post is missing from the Helix 2.5 card entirely.
What Figure measured
Figure trained four models on nested subsets of Index spanning an eightfold increase in data, holding model size and downstream training fixed. Held-out action-prediction loss “fell predictably with each doubling of Index”.
The forecast
Using only the smaller runs, Figure says it predicted its largest run’s test loss “to four decimal places before training began”, with a forecasting error of 0.54 per cent of the variation across the full range. It calls this “the first human-to-robot transfer scaling law measured on a humanoid”.
Why it matters more than 56 per cent
A success rate describes one model. A scaling law describes what the next doubling of data should buy. If it holds, Figure can plan data collection and compute the way language model labs do, and the $3.5 billion Nscale commitment becomes a forecastable investment rather than a gamble.
The limits Figure states
It measures data scaling only, with model size and downstream training fixed, and it predicts loss rather than task success. Figure’s own wording is that it “suggests” language-model scaling “may extend” to humanoids. That is appropriately cautious, and worth repeating whenever the scaling law is quoted.
The Badge: #4 in Robots and Devices
Every card in this series gets its rank badge walked against the category it links to. Past cards have produced absent entries, predecessors in the successor’s slot and duplicate ordinals. Helix 2.5’s badge is the first to raise a different question: not whether the number is right, but what is being ranked.
The number is exact
The badge links to Robots and Devices, which runs to five pages (“Page 2 of 5”). Helix 2.5 sits at position 4 on page 1, exactly where the badge says.
| Card | Badge | Position on category page | Votes behind its rating |
|---|---|---|---|
| Chessnut | #1 | 1 | 12 |
| Tesla Optimus | #2 | 2 | 2 |
| NEO by 1X Technologies | #3 | 3 | 2 |
| Helix 2.5 | #4 | 4 | 2 |
| Figure 03 | #5 | 5 | 6 |
| Figure AI | #11 | 11 | 3 |
The controls are clean
All five controls we opened matched their slots exactly. The badge is a working ordinal in this category, so #4 is a real position, not decoration.
The oddity is the ranking itself
Helix 2.5 is software. It now outranks Figure 03, the robot it runs on, and Figure AI, the company’s own card, eleven places down. A category called Robots and Devices therefore lists the same company three times, and puts its model above its hardware on the strength of a card that is less than a day old.
What the rating hides
The 4.5 rating rests on two votes, the same as Tesla Optimus and NEO. Chessnut, the chess board at #1, has twelve. None of these ratings says much, and the rank order clearly reflects something other than votes.
The Alternatives Strip and the Ranked Lists
Two more checks from the series routine, both of which show how a directory positions a research model.
The alternatives
The card lists eight alternatives: Chessnut, Apolosign, Figure 03, NEO, PettiChat, Figure AI, Unitree G1 and Microduck. Mapped to the category, they are positions 1, 14, 5, 3, 12, 11, 8 and 7. The strip is the top of the category, not a list of rival models. It includes a smart chess board, a digital family calendar and a pet-communication device, and it omits the one comparable model the card’s own review mentions indirectly: Physical Intelligence’s work on home generalisation.
Propagation on day one
At about 12:55 UTC we counted mentions on AIxploria’s Top 100, Ultimate List and Free AI pages. Helix 2.5 scored zero on all three, which is normal for a same-day card. The word “Helix” appeared twice on the Top 100, both inside the tooltip of the Figure 03 entry, whose description mentions “integrated AI (Helix)”.
What that says about browsing
A reader who finds Figure through the Top 100 lands on the Figure 03 hardware card, whose description predates Helix 2.5. The new model is visible only to readers who browse the category or search for it. For a directory, the card and its lists run on different clocks.
Helix 2.5 Against the Rest of the Field
The card mentions 1X’s NEO and Tesla’s Optimus as “two different roads”. Coverage of the launch added more useful comparisons, with caveats that matter.
Sunday Robotics
Startup Fortune reports that Sunday Robotics recently claimed a 99.1 per cent success rate, 778 of 785 attempts, for its ACT-2 system folding laundry in unseen environments, citing Humanoids Daily. The report itself cautions that the results “aren’t directly comparable”: different tasks, rules and definitions of unseen.
Physical Intelligence
techAU’s coverage set Helix 2.5 beside Physical Intelligence’s pi 0.5, a vision-language-action model trained on a mix of web data and robot data from several platforms, which it reports approached in-home baselines after training across roughly 100 environments. It is a different recipe: broad multimodal data rather than human video.
The fair comparison
The honest summary is that nobody yet runs a shared benchmark for household humanoids. Figure’s contribution is not the highest number but a strict protocol: blind evaluation, fixed checkpoints, timeouts and safety interventions counted as failures. Our earlier report on a humanoid robot trained on human motion data found the same pattern: the protocol tells you more than the headline.
What Helix 2.5 Means Outside the Lab
For businesses following robotics, the release changes some expectations and leaves others exactly where they were.
What changed
The argument that every deployment needs site-specific data collection is weaker today. If a model pretrained on human video can make a stranger’s bed two times in three, the cost of rolling robots into new sites could fall sharply. That is the claim investors will focus on.
What did not
A 44 per cent failure rate across whole tasks means a robot that needs supervision. The homes were empty rentals, and techAU notes engineers visible “within arm reach” in some clips. Nothing is for sale, and nothing about safety around children, pets or moving people was tested.
Where to watch
The things to watch are Figure’s promised pilot deployments, any independent evaluation, and whether the scaling law holds at the next doubling. For more on physical AI and model releases, see our AI models, tools and releases hub and our report on compact training for crowd navigation.
Helix 2.5 on AIxploria: The Verdict
As a summary of Figure’s research, the Helix 2.5 card is good. Its core numbers match the announcement, its caveats are fair, and its FAQ explains zero-shot more carefully than much of the launch coverage. It also leaves out two of the most important details: the 1.90 per cent cap on any task in the pretraining data, and the scaling law.
The weak points are the labels around it. A “Paid” tag on a model that cannot be bought, an unsourced sub-$20,000 price, and a rank badge that is numerically exact but places software above its own robot. None of these is a research error. All of them are what a browsing reader will absorb first.
For readers following the AIxploria series, this is the fifth badge in a row that proved to be a working ordinal once the controls were checked. It follows the MAI-Transcribe-2 card, whose badge duplicated another tool’s number. Helix 2.5 is the first case where the number is right and the category is the question.
References
AIxploria Robots and Devices category
Helix 2.5: Zero-Shot 30-Home Generalization (Figure)
Figure and Nscale Sign Strategic Partnership (Figure)
Introducing Helix 02: Full-Body Autonomy (Figure)
Introducing Figure 03 (Figure)
Figure AI’s Helix 2.5 Robot Made Beds in 30 Homes It Had Never Seen (Startup Fortune)
Figure releases Helix 2.5, boosting robot task success rate to 56% (Crypto Briefing)
Figure Unveils Helix 2.5 With Zero-Shot Humanoid Generalization (The AI Insider)
Figure reveals Helix 2.5 AI model as humanoid robots tackle chores (techAU via MSN)
Zero-shot learning (Wikipedia)
Humanoid Robot Learns to Sprint and Perform Spin Kicks Using AI Trained on Human Motion Data
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.