Published on

Beyond the Viral Demo: Sunday Robotics Claims 99.1% Zero-Shot Success in Laundry Folding with ACT-2

Humanoids Daily
Written byHumanoids Daily
  • Sunday Robotics introduced the ACT-2 model, achieving a 99.1% autonomous success rate in folding laundry across unseen home environments.
  • The company is proposing a new industry standard called a "Solve," which evaluates tasks based on performance, scope, and adaptation cost rather than isolated demonstrations.
  • By scaling pretraining on diverse data, ACT-2 closes the "generalization gap," allowing the model to learn new, transferable behaviors from a single fine-tuning example.
  • The system exhibited emergent behaviors, such as recovering dropped garments and continuing tasks despite physical disturbances from humans or lighting changes.
  • The breakthrough relies on the wheeled, telescoping design of the Memo robot, which provides the physical range necessary for handling varying garment sizes.

In 1903, the Wright brothers’ first flight lasted twelve seconds—proving powered flight was possible long before it became reliable. Sunday Robotics applies this analogy to the current state of humanoid robotics, arguing that the industry must move past early capability demonstrations and begin measuring how performance holds up as real-world conditions change.

Having previously stated their intent to move beyond the "viral demo", Sunday has now introduced ACT-2, a foundation model designed to tackle the long-tail edge cases of household tasks. The company has chosen to start with what it calls the "ultimate challenge": laundry folding.

In a newly released video (watch it below) and technical blog post, the company claims a 99.1% zero-shot success rate for laundry folding across diverse, unseen environments. Crucially, this performance required zero per-home data collection, zero expert demonstrations in target homes, and zero post-training adaptation.

A grid of 25 video stills showing Sunday Memo robots folding laundry in diverse, unseen bedroom environments.
To prove zero-shot generalization, Sunday Robotics evaluated ACT-2 across 25 unseen real-world home environments. The Memo robot achieved a 99.1% success rate across varied bed configurations, sheet colors, and lighting conditions without any home-specific fine-tuning. Image: Sunday Robotics

Closing the "Generalization Gap"

The central technical unlock of ACT-2 builds upon the foundation of ACT-1, focusing heavily on solving the "generalization gap". Historically in robotics, training on narrow, curated data improves performance locally but fails in novel environments—a phenomenon known as overfitting.

Sunday discovered that by scaling pretraining with high-diversity data—collected via their proprietary Skill Capture Glove ecosystem—the generalization gap shrinks dramatically. As the pretrained model grows stronger, local performance gains from rapid, in-house post-training iterations become generalizable to unseen real-world homes. In fact, Sunday demonstrated that ACT-2 can learn a new, transferable folding behavior from a single supervised fine-tuning demonstration.

Defining a "Solve"

To contextualize this milestone, Sunday is pushing a new evaluation framework called a "Solve". A Solve requires developers to explicitly define three parameters: performance, scope, and adaptation cost.

This framework serves as a direct critique of the broader industry's reliance on highly controlled demonstrations—a dynamic Sunday has frequently clashed over, most notably during the recent "dishwasher wars" with Figure. For ACT-2's laundry Solve, the stated scope encompassed nine distinct garment types (from XXS baby clothes to 8XL shirts) in arbitrarily crumpled starting configurations. The adaptation cost was defined as exactly zero per home.

An overhead wide shot of Sunday Robotics' testing facility featuring several Memo robots folding laundry on beds simultaneously.
Multiple Memo robots perform autonomous laundry folding inside Sunday's lab facility. This high-throughput post-training setup enables Sunday to rapidly hill-climb model reliability before transferring gains to unseen deployment environments. Image: Sunday Robotics

Quality and Emergent Behaviors

Performance was measured across 785 autonomous attempts, graded by independent annotators on a strict five-star rubric that penalized overfolding, misalignment, and stack instability. The 778 successful folds achieved an average quality score of 4.72 out of 5 stars, with 98.3% meeting the four- or five-star quality bar. Folds were completed in a median time of 2 minutes and 13 seconds.

Beyond strict metrics, ACT-2 displayed impressive emergent behaviors under long-tail conditions. Because real homes are chaotic, the model learned to treat messy intermediate states as recoverable rather than dead ends. The Memo robot successfully retrieved clothes dropped on the floor, handled severe crumpling, and maintained its workflow despite adversarial perturbations, such as a child playing with Legos in its workspace or sudden shifts from dark to bright lighting.

Handling objects as small as a baby shirt and as large as a bath towel requires significant physical range. Sunday notes this validates their full-stack approach: Memo utilizes a wheeled base and telescoping spine to reposition and lean into workspaces, extending its operating range far beyond that of a fixed tabletop rig.

Sunday Robotics backed this research with a recent $165 million Series B meant to propel them out of the lab. As they prepare to deploy Memo to families in their 2026 Beta Program, ACT-2 serves as a formidable proof point that their data-first, generalization-heavy recipe could successfully scale to other household tasks like vacuuming, zipping, and toy organization.

Comments

No comments yet. Be the first to share your thoughts!

Share this article

Stay Ahead in Humanoid Robotics

Get the latest developments, breakthroughs, and insights in humanoid robotics — delivered straight to your inbox.