- Published on
Dyna-2 Proves Scaling Laws for Robotics: 1 Million Hours of Human Video Unlocks Zero-Shot Dexterity
One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

- Dyna Robotics has unveiled Dyna-2, a world-action model (WAM) pre-trained on over one million hours of human video data.
- The model demonstrates a first-of-its-kind human-to-robot transfer scaling law, proving that scaling human video data predictably improves zero-shot performance on unseen robot hardware.
- Dyna-2's architecture relies on video co-training (predicting future video states), which the company claims is essential for cross-embodiment generalization, directly challenging the industry standard of Vision-Language-Action (VLA) models.
- In real-world zero-shot deployments at customer sites, Dyna-2 achieved an 87% pass rate, vastly outperforming the 46% pass rate of its VLA predecessor, Dyna-1.
- The release also introduces a novel one-step video generation distillation pipeline, dropping inference latency by two orders of magnitude for downstream planning.
The robotics industry has spent the last year fiercely debating the architecture of physical intelligence, caught between fine-tuning existing language models and building native "world models" from scratch. Today, Dyna Robotics delivered what may be the strongest empirical evidence yet for the latter, unveiling Dyna-2—a world-action model (WAM) pre-trained on a staggering one million hours of human video data.
According to the company's technical report and an accompanying social media thread, this massive scale has unlocked a holy grail of embodied AI: a human-to-robot transfer scaling law. In short, Dyna-2 proves that feeding a model more human video predictably improves its ability to control a robot it has never seen before.
Bridging the Embodiment Gap
The historical bottleneck in robot learning has been the data itself. While teleoperation yields high-quality, action-labeled data, it is slow and expensive to collect. The theoretical alternative is to learn from the boundless supply of human video on the internet, but translating a human hand's movement into a robotic gripper's action—the "embodiment gap"—has proven exceedingly difficult.
Dyna-2 attacks this problem purely through scale and objective design. The company curated nested subsets of egocentric human manipulation videos, scaling from 1,000 to 1,000,000 hours, keeping proportions from each source identical. When evaluated zero-shot on 39 distinct robot tasks across two stationary, bimanual platforms, Dyna-2's performance improved monotonically as the human pre-training data increased. An inflection point emerged between 10,000 and 100,000 hours, suggesting that cross-embodiment knowledge transfer emerges naturally if the model simply sees enough human activity.
Stay ahead in humanoid robotics
One email a week: the launches, funding and research that mattered. Join the researchers, founders and technologists following the field.
Furthermore, this zero-shot capability extended to post-training. With just a few hours of robot-specific data and zero human-robot alignment, post-trained Dyna-2 models solved tasks ranging from manipulating deformable objects to untwisting bottle caps.
Video as the New Scaling Axis
Crucially, Dyna Robotics found that predicting next actions alone is insufficient to bridge this gap. In an ecosystem currently dominated by Vision-Language-Action (VLA) models, Dyna-2 relies heavily on "video co-training"—forcing the model to predict both future actions and future video frames simultaneously. This aligns closely with the cognitive architecture bets being placed by researchers like Yann LeCun, who has dismissed LLM-based approaches as fundamentally flawed for physical reasoning.
Dyna's controlled ablations show that joint denoising unanimously beats action-only training at every scale on all 39 tasks. Furthermore, the company established that video itself is a "new scaling axis". Even when action-labeled data is capped, simply scaling unannotated human video continues to improve the model's generalization capabilities across different robotic embodiments. This empirical backing provides a critical validation for companies like 1X, which recently pivoted to bet everything on world models to digest raw internet video.
Beating VLAs and Surviving the Real World
To prove the WAM architecture's superiority, Dyna Robotics ran an "apple-to-apple" comparison between an early version of Dyna-2 and their production VLA model, Dyna-1. Despite using datasets and hyperparameters tuned specifically for the VLA, the WAM achieved 1.55x the success rate of its predecessor. Qualitative tests showcased Dyna-2 completing tasks under extreme conditions, such as flickering "disco lights," complete darkness, and active human interference where a researcher undid the robot's work in real-time.
This robustness translates directly to commercial viability. In zero-shot deployments at novel customer sites, Dyna-2 achieved an 87% production pass rate, a massive leap over Dyna-1's 46%.
The new architecture also exhibited emergent language-following capabilities, an area where continuous action models often struggle without destroying pre-trained representations. Video co-training boosted language-following success rates from 35% to 67%, and scaling the corpus pushed it to 96% on internal benchmarks.
Speeding Up the "Dream"
A known drawback of generative world models is the "reactivity gap"—the immense compute required to predict future states before a robot can physically act. To address this, Dyna-2 introduces a novel one-step video generation distillation pipeline. By matching a student model against an evolving target distribution rather than a fixed one, the company reduced video generation latency from 10,203 milliseconds down to just 110 milliseconds on a single H100 GPU. This allows the model to "dream" high-fidelity, instruction-conditioned futures almost instantly without sacrificing quality.
With rivals like Generalist AI raising hundreds of millions to train physical models from scratch, and others building massive simulation engines to evaluate them, Dyna-2's million-hour milestone marks a significant escalation in the physical AI arms race. As the company states, one million hours is "only the beginning of a new era of scaling for robotics".
Share this article
Read next
- Published on
- Reading time
- 5 min read
The "OpenAI of Robotics" Debate Reignites as GPT-6 Astra Paints in Real Life
- Published on
- Reading time
- 5 min read
Nucleus Pivots to Wheeled Base for "Nucleus II" Following Factory Floor Pushback
- Published on
- Reading time
- 6 min read
Inside XPENG’s Bet on IRON: He Xiaopeng on Why the Hardest Path in Humanoid Robotics Is Never Crowded
- Published on
- Reading time
- 4 min read
Figure AI Reports Rapid Growth for Index, Surpassing 69,000 Weekly Active Users
- Published on
- Reading time
- 10 min read
The Policy Is a Video: How Markov Robotics Runs Dexterous Manipulation on LTX-2.5
- Published on
- Reading time
- 5 min read
Dynamic Creatures Emerges From Stealth With Boston Dynamics Backing to Build Expressive Guest-Facing Robots
Stay ahead in humanoid robotics
One email a week: the launches, funding and research that mattered. Join the researchers, founders and technologists following the field.













