Helix 2.5, Figure’s new AI model for humanoid robots, has landed on AIxploria. The directory published its card at 04:59 UTC on 18 September 2026, about sixteen hours after Figure announced the model, with a 4.5 rating from two votes, a “Verified Tool” label, a “Paid” tag and a badge reading “#4 in Robots and Devices”. By early afternoon it had 7,398 views and 63 upvotes.

The model itself is a genuine research result. Figure took one humanoid into 30 Bay Area homes it had never seen, collected no data in any of them, and had it tidy toys, fold towels and make beds. With its new pretraining the robot succeeded 56 per cent of the time; the same system trained from scratch managed 9 per cent.

We checked the card the way this series checks every listing: claim by claim against Figure’s own announcement and related pages, with the badge walked against the category it links to and the ranked lists searched for the new entry. The numbers mostly hold. The labels are another matter: a model nobody can buy is tagged “Paid”, and a piece of software now ranks above the robot it runs on.

What the Helix 2.5 Card Says

helix 2 5 figure humanoid robot aixploria b helix 2.5 oval laundry hamper with two handle slots

The card is a full editorial review headed “one model, three chores, and 30 homes the robot had never entered”. It is filed in two categories, Robots and Devices and Latest AI.

The summary

The blurb calls Helix 2.5 “a humanoid robot controlled by a neural network that tidies up, folds laundry, and makes beds in homes it has never visited (without collecting data on site)”. The review adds the headline numbers: 56 per cent success over 420 attempts, against 9 per cent without pretraining, and “nothing is for sale at this point, neither the model nor the robot.”

Pros and cons

Four pros: zero-shot results measured in 30 real homes, one network for walking, seeing and grasping, recovery after failed grabs, and half the task data needed by Helix 02. Three cons: 56 per cent is “still far from daily reliability”, there are no weights, API or product, and it is a research-stage system that is “watch-only for now”.

The verdict

The review ends that a robot making a stranger’s bed roughly half the time “may sound modest, yet it is the strongest sign so far that a chore learned once can travel to new places.” That is a fair reading of Figure’s evidence, and it is a more careful verdict than most launch coverage managed.

Checking the Helix 2.5 Card Against Figure's Announcement

helix 2 5 figure humanoid robot aixploria c single bed frame with a thick mattress and pillow

Figure published “Helix 2.5: Zero-Shot 30-Home Generalization” on 17 September 2026. We compared every checkable claim on the card with that post, Figure’s earlier Index and Nscale announcements, and the news reports that carried additional figures.

Card claimWhat Figure or coverage saysVerdict
30 homes, no data collected on site“30 Bay Area homes with zero data collected in any of them”Matches
56% success vs 9% without Index“from 9% to 56%”, with all other variables held fixedMatches
237 of 420 trials completedNot in the post’s text; reported by Startup Fortune and Crypto Briefing with a per-task splitConsistent with coverage
13 to 15 toys per tidy“All 13-15 toys scattered in the scene”Matches
Half the task data of Helix 02“half as much task-specific data as a representative Helix 02 behavior”Matches
Pretrained from scratch on Index“pretrained from random initialization entirely on Index”Matches
35 minutes of human data per second“roughly 35 minutes of new human experience every second”Matches
$3.5 billion of compute committedNscale deal: “initial commitment of $3.5 billion”, intent to scale past $6 billionMatches, context missing
Homes were rentedBrett Adcock: “We rented 30 homes in the Bay Area”Matches the CEO’s post
Runs on the Figure 03 robotFigure’s post does not name the robot; Crypto Briefing says Figure 03Not stated by Figure
Consumer price below $20,000Not on any Figure page we readUnsourced
At least one success in every homeNot in the post’s textNot found
“Paid” tagThe card’s own review: “Nothing is for sale”Self-contradictory

The score

Of thirteen claims, eight match Figure or its chief executive, one is consistent with independent coverage, two could not be found, one is unsourced and one contradicts the card’s own text. On the research itself, the Helix 2.5 card is accurate. The drift is all in the commercial details.

The price line

The card says a consumer price “below $20,000” is targeted, and hedges that “nothing official has been confirmed yet”. We found no such figure on Figure’s Helix 2.5 post, the Figure 03 launch page or the Index and Nscale announcements. With that hedge the line is honest, but a reader skimming the card will remember the number, not the caveat.

The "Paid" Tag on a Model Nobody Can Buy

helix 2 5 figure humanoid robot aixploria d heap of five toy bricks with round studs

Every AIxploria card carries a pricing tag, and it is drawn from the structured data behind the page. For Helix 2.5 that data says “Paid”.

What the card itself says

The same page states that “no model weights, API or pricing have been published for Helix 2.5”, that the robot “cannot be bought by the public”, and, in its FAQ, that “Figure sells nothing to the public and has opened no pre-orders”. The only available step, it says, is registering interest on Figure’s website.

Why the tag is wrong in a specific way

“Paid” implies a transaction is possible. “Free” would be equally wrong. The honest label is something like “Not available”, which AIxploria’s tag set does not appear to include. This is the same class of error the series found on the MAI-Transcribe-2 card, whose “Free” tag sat above an FAQ quoting a price.

Why it matters to readers

Directory tags drive filters. A reader browsing paid robotics tools will find Helix 2.5 beside products with checkout pages. The review text fixes the impression for anyone who reads it; the tag misleads anyone who does not.

The Results Behind Helix 2.5, Task by Task

helix 2 5 figure humanoid robot aixploria e clothes iron standing on its flat heel

Figure’s post gives the headline 56 per cent and a chart. The per-task numbers come from reports that read the trial counts: Startup Fortune and Crypto Briefing both give 237 successes out of 420 attempts, split across three tasks of 140 trials each.

TaskSuccessesAttemptsSuccess rateWhat counted as success
Bed making9414067%Both pillows and comforter corners at the top, comforter pulled smooth
Towel folding8714062%All towels folded and placed in the basket
Living room tidy5614040%All 13 to 15 toys picked and placed in the basket
All three23742056.4%Full completion only; no partial credit

The arithmetic checks out

94 plus 87 plus 56 is 237, and 237 divided by 420 is 56.4 per cent, which rounds to Figure’s 56. Three tasks of 140 trials across 30 homes is about 4.7 trials per task per home, so each home saw roughly fourteen attempts in total.

Helix 2.5 zero-shot success by task, 30 unseen homes (trial counts as reported)
Bed making, 94 of 140 67%
Towel folding, 87 of 140 62%
All tasks, 237 of 420 56%
Living room tidy, 56 of 140 40%
Same policy trained from scratch, all tasks 9%

Why toys are hardest

The tidy task needs every one of 13 to 15 toys placed, with a one-minute timeout per toy. One missed toy fails the whole trial. Bed making, by contrast, involves a handful of large objects. The spread tells you more about how strictly each task was graded than about which chore is intrinsically harder.

The six-fold claim

Figure describes 56 per cent as “over 6x higher” than 9 per cent. That is correct: 56 divided by 9 is 6.2. Because only the pretraining changed between the two policies, Figure argues the gap “directly measures its contribution”.

How Strict Was the Helix 2.5 Grading?

helix 2 5 figure humanoid robot aixploria f trigger spray bottle with a flat head

The card mentions “no partial credit”. Figure’s appendix goes further, and it is the part of the release that makes the numbers credible.

Timeouts on every object

Each toy got a one-minute timeout, and exceeding it aborted the rollout. Each towel got three minutes. Each pillow and each side of the comforter got one minute. Graders also recorded attempts per toy and graded the quality of each fold, even when the trial failed.

Safety interventions count as failures

“If a human intervention is necessary for safety, that rollout is aborted and failed.” That closes the most common loophole in robot demos, where a hand quietly steadies the robot off camera.

Fixed checkpoints, blind evaluation

Each task used “a single fixed checkpoint across all 30 homes”, no weights were adapted to the homes, and no evaluation data was used to pick checkpoints. Evaluation objects were set aside in advance and checked by an AI model and then a human to confirm they were absent from the training data.

What “zero-shot” does and does not mean

Figure is careful here, and so is the card’s FAQ. The homes and objects were new. The behaviours were not: they were taught through fine-tuning data collected elsewhere. “Zero-shot” covers places and objects, not the skills themselves.

Index: The Dataset Doing the Work

Most of what is new in Helix 2.5 comes from Index, the human-video dataset Figure unveiled three weeks earlier. The card says Index “does the heavy lifting”, and Figure’s own numbers support that.

What Index is

Figure brought Index out of stealth on 25 August 2026 as “the most diverse robot training dataset ever built”. People record themselves doing real tasks through a phone app; Figure calls them Creators and says it had paid them $15 million by launch.

The scale, then and now

At launch Figure reported 264,000 app downloads across 108 countries, over 44,000 weekly active users, more than 16 million uploaded videos and “30 minutes of video uploads every second”. By the Nscale announcement on 3 September the rate was 35 minutes a second, and the Helix 2.5 post repeats “roughly 35 minutes”.

Index upload rate, minutes of human video per second (Figure announcements)
25 Aug 2026, Index launch 30 min/s
3 Sep 2026, Nscale partnership 35 min/s
17 Sep 2026, Helix 2.5 post about 35 min/s

Thirty-five minutes a second is 35 times 86,400 seconds, or about 3.0 million minutes of video a day, roughly 5.8 years of footage every 24 hours. Figure’s own launch figure, “4.9 years of human work every day”, was set at the 30-minute rate, and 4.9 scaled by 35/30 gives 5.7, so the two statements agree.

DateFigure announcementWhat it added
9 Oct 2025Introducing Figure 03Palm cameras, fingertip tactile sensors, a home-safe design and the BotQ factory
27 Jan 2026Introducing Helix 02Whole-body control; a 4-minute autonomous dishwasher task
25 Aug 2026Introducing IndexHuman-video dataset; 30 minutes uploaded per second
3 Sep 2026Nscale partnershipUp to 100,000 Vera Rubin GPUs; $3.5bn initial compute commitment
17 Sep 2026Helix 2.5Index-pretrained model; 30-home zero-shot evaluation
18 Sep 2026AIxploria cardPublished 04:59 UTC, about 16 hours after Figure’s post

How broad the pretraining is

Figure says “no single evaluation task makes up more than 1.90% of the Index pretraining dataset.” That is an important detail the card leaves out, because it is the answer to the obvious objection that the model simply saw lots of bed making.

A different starting point from Helix 02

Helix 02 began from a pretrained vision-language model and used a whole-body controller, System 0, trained on over 1,000 hours of human motion data and sim-to-real reinforcement learning. Helix 2.5 was pretrained from random weights on Index alone. That is a bet that human video can replace internet-scale vision-language pretraining for a robot’s foundation.

The Scaling Law the Card Leaves Out

The most consequential claim in Figure’s post is missing from the Helix 2.5 card entirely.

What Figure measured

Figure trained four models on nested subsets of Index spanning an eightfold increase in data, holding model size and downstream training fixed. Held-out action-prediction loss “fell predictably with each doubling of Index”.

The forecast

Using only the smaller runs, Figure says it predicted its largest run’s test loss “to four decimal places before training began”, with a forecasting error of 0.54 per cent of the variation across the full range. It calls this “the first human-to-robot transfer scaling law measured on a humanoid”.

Why it matters more than 56 per cent

A success rate describes one model. A scaling law describes what the next doubling of data should buy. If it holds, Figure can plan data collection and compute the way language model labs do, and the $3.5 billion Nscale commitment becomes a forecastable investment rather than a gamble.

The limits Figure states

It measures data scaling only, with model size and downstream training fixed, and it predicts loss rather than task success. Figure’s own wording is that it “suggests” language-model scaling “may extend” to humanoids. That is appropriately cautious, and worth repeating whenever the scaling law is quoted.

The Badge: #4 in Robots and Devices

Every card in this series gets its rank badge walked against the category it links to. Past cards have produced absent entries, predecessors in the successor’s slot and duplicate ordinals. Helix 2.5’s badge is the first to raise a different question: not whether the number is right, but what is being ranked.

The number is exact

The badge links to Robots and Devices, which runs to five pages (“Page 2 of 5”). Helix 2.5 sits at position 4 on page 1, exactly where the badge says.

CardBadgePosition on category pageVotes behind its rating
Chessnut#1112
Tesla Optimus#222
NEO by 1X Technologies#332
Helix 2.5#442
Figure 03#556
Figure AI#11113

The controls are clean

All five controls we opened matched their slots exactly. The badge is a working ordinal in this category, so #4 is a real position, not decoration.

The oddity is the ranking itself

Helix 2.5 is software. It now outranks Figure 03, the robot it runs on, and Figure AI, the company’s own card, eleven places down. A category called Robots and Devices therefore lists the same company three times, and puts its model above its hardware on the strength of a card that is less than a day old.

What the rating hides

The 4.5 rating rests on two votes, the same as Tesla Optimus and NEO. Chessnut, the chess board at #1, has twelve. None of these ratings says much, and the rank order clearly reflects something other than votes.

The Alternatives Strip and the Ranked Lists

Two more checks from the series routine, both of which show how a directory positions a research model.

The alternatives

The card lists eight alternatives: Chessnut, Apolosign, Figure 03, NEO, PettiChat, Figure AI, Unitree G1 and Microduck. Mapped to the category, they are positions 1, 14, 5, 3, 12, 11, 8 and 7. The strip is the top of the category, not a list of rival models. It includes a smart chess board, a digital family calendar and a pet-communication device, and it omits the one comparable model the card’s own review mentions indirectly: Physical Intelligence’s work on home generalisation.

Propagation on day one

At about 12:55 UTC we counted mentions on AIxploria’s Top 100, Ultimate List and Free AI pages. Helix 2.5 scored zero on all three, which is normal for a same-day card. The word “Helix” appeared twice on the Top 100, both inside the tooltip of the Figure 03 entry, whose description mentions “integrated AI (Helix)”.

What that says about browsing

A reader who finds Figure through the Top 100 lands on the Figure 03 hardware card, whose description predates Helix 2.5. The new model is visible only to readers who browse the category or search for it. For a directory, the card and its lists run on different clocks.

Helix 2.5 Against the Rest of the Field

The card mentions 1X’s NEO and Tesla’s Optimus as “two different roads”. Coverage of the launch added more useful comparisons, with caveats that matter.

Sunday Robotics

Startup Fortune reports that Sunday Robotics recently claimed a 99.1 per cent success rate, 778 of 785 attempts, for its ACT-2 system folding laundry in unseen environments, citing Humanoids Daily. The report itself cautions that the results “aren’t directly comparable”: different tasks, rules and definitions of unseen.

Physical Intelligence

techAU’s coverage set Helix 2.5 beside Physical Intelligence’s pi 0.5, a vision-language-action model trained on a mix of web data and robot data from several platforms, which it reports approached in-home baselines after training across roughly 100 environments. It is a different recipe: broad multimodal data rather than human video.

The fair comparison

The honest summary is that nobody yet runs a shared benchmark for household humanoids. Figure’s contribution is not the highest number but a strict protocol: blind evaluation, fixed checkpoints, timeouts and safety interventions counted as failures. Our earlier report on a humanoid robot trained on human motion data found the same pattern: the protocol tells you more than the headline.

What Helix 2.5 Means Outside the Lab

For businesses following robotics, the release changes some expectations and leaves others exactly where they were.

What changed

The argument that every deployment needs site-specific data collection is weaker today. If a model pretrained on human video can make a stranger’s bed two times in three, the cost of rolling robots into new sites could fall sharply. That is the claim investors will focus on.

What did not

A 44 per cent failure rate across whole tasks means a robot that needs supervision. The homes were empty rentals, and techAU notes engineers visible “within arm reach” in some clips. Nothing is for sale, and nothing about safety around children, pets or moving people was tested.

Where to watch

The things to watch are Figure’s promised pilot deployments, any independent evaluation, and whether the scaling law holds at the next doubling. For more on physical AI and model releases, see our AI models, tools and releases hub and our report on compact training for crowd navigation.

Helix 2.5 on AIxploria: The Verdict

As a summary of Figure’s research, the Helix 2.5 card is good. Its core numbers match the announcement, its caveats are fair, and its FAQ explains zero-shot more carefully than much of the launch coverage. It also leaves out two of the most important details: the 1.90 per cent cap on any task in the pretraining data, and the scaling law.

The weak points are the labels around it. A “Paid” tag on a model that cannot be bought, an unsourced sub-$20,000 price, and a rank badge that is numerically exact but places software above its own robot. None of these is a research error. All of them are what a browsing reader will absorb first.

For readers following the AIxploria series, this is the fifth badge in a row that proved to be a working ordinal once the controls were checked. It follows the MAI-Transcribe-2 card, whose badge duplicated another tool’s number. Helix 2.5 is the first case where the number is right and the category is the question.

References