- Published on
Dyna-2 Proves Scaling Laws for Robotics: 1 Million Hours of Human Video Unlocks Zero-Shot Dexterity

- Dyna Robotics has unveiled Dyna-2, a world-action model (WAM) pre-trained on over one million hours of human video data.
- The model demonstrates a first-of-its-kind human-to-robot transfer scaling law, proving that scaling human video data predictably improves zero-shot performance on unseen robot hardware.
- Dyna-2's architecture relies on video co-training (predicting future video states), which the company claims is essential for cross-embodiment generalization, directly challenging the industry standard of Vision-Language-Action (VLA) models.
- In real-world zero-shot deployments at customer sites, Dyna-2 achieved an 87% pass rate, vastly outperforming the 46% pass rate of its VLA predecessor, Dyna-1.
- The release also introduces a novel one-step video generation distillation pipeline, dropping inference latency by two orders of magnitude for downstream planning.
The robotics industry has spent the last year fiercely debating the architecture of physical intelligence, caught between fine-tuning existing language models and building native "world models" from scratch. Today, Dyna Robotics delivered what may be the strongest empirical evidence yet for the latter, unveiling Dyna-2—a world-action model (WAM) pre-trained on a staggering one million hours of human video data.
According to the company's technical report and an accompanying social media thread, this massive scale has unlocked a holy grail of embodied AI: a human-to-robot transfer scaling law. In short, Dyna-2 proves that feeding a model more human video predictably improves its ability to control a robot it has never seen before.
Bridging the Embodiment Gap
The historical bottleneck in robot learning has been the data itself. While teleoperation yields high-quality, action-labeled data, it is slow and expensive to collect. The theoretical alternative is to learn from the boundless supply of human video on the internet, but translating a human hand's movement into a robotic gripper's action—the "embodiment gap"—has proven exceedingly difficult.
Dyna-2 attacks this problem purely through scale and objective design. The company curated nested subsets of egocentric human manipulation videos, scaling from 1,000 to 1,000,000 hours, keeping proportions from each source identical. When evaluated zero-shot on 39 distinct robot tasks across two stationary, bimanual platforms, Dyna-2's performance improved monotonically as the human pre-training data increased. An inflection point emerged between 10,000 and 100,000 hours, suggesting that cross-embodiment knowledge transfer emerges naturally if the model simply sees enough human activity.
Furthermore, this zero-shot capability extended to post-training. With just a few hours of robot-specific data and zero human-robot alignment, post-trained Dyna-2 models solved tasks ranging from manipulating deformable objects to untwisting bottle caps.
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws: • world-action models exhibit scaling law on human data across four orders of magnitude, from 1000
Video as the New Scaling Axis
Crucially, Dyna Robotics found that predicting next actions alone is insufficient to bridge this gap. In an ecosystem currently dominated by Vision-Language-Action (VLA) models, Dyna-2 relies heavily on "video co-training"—forcing the model to predict both future actions and future video frames simultaneously. This aligns closely with the cognitive architecture bets being placed by researchers like Yann LeCun, who has dismissed LLM-based approaches as fundamentally flawed for physical reasoning.
Dyna's controlled ablations show that joint denoising unanimously beats action-only training at every scale on all 39 tasks. Furthermore, the company established that video itself is a "new scaling axis". Even when action-labeled data is capped, simply scaling unannotated human video continues to improve the model's generalization capabilities across different robotic embodiments. This empirical backing provides a critical validation for companies like 1X, which recently pivoted to bet everything on world models to digest raw internet video.
Beating VLAs and Surviving the Real World
To prove the WAM architecture's superiority, Dyna Robotics ran an "apple-to-apple" comparison between an early version of Dyna-2 and their production VLA model, Dyna-1. Despite using datasets and hyperparameters tuned specifically for the VLA, the WAM achieved 1.55x the success rate of its predecessor. Qualitative tests showcased Dyna-2 completing tasks under extreme conditions, such as flickering "disco lights," complete darkness, and active human interference where a researcher undid the robot's work in real-time.
This robustness translates directly to commercial viability. In zero-shot deployments at novel customer sites, Dyna-2 achieved an 87% production pass rate, a massive leap over Dyna-1's 46%.
The new architecture also exhibited emergent language-following capabilities, an area where continuous action models often struggle without destroying pre-trained representations. Video co-training boosted language-following success rates from 35% to 67%, and scaling the corpus pushed it to 96% on internal benchmarks.
Speeding Up the "Dream"
A known drawback of generative world models is the "reactivity gap"—the immense compute required to predict future states before a robot can physically act. To address this, Dyna-2 introduces a novel one-step video generation distillation pipeline. By matching a student model against an evolving target distribution rather than a fixed one, the company reduced video generation latency from 10,203 milliseconds down to just 110 milliseconds on a single H100 GPU. This allows the model to "dream" high-fidelity, instruction-conditioned futures almost instantly without sacrificing quality.
With rivals like Generalist AI raising hundreds of millions to train physical models from scratch, and others building massive simulation engines to evaluate them, Dyna-2's million-hour milestone marks a significant escalation in the physical AI arms race. As the company states, one million hours is "only the beginning of a new era of scaling for robotics".
Share this article
Stay Ahead in Humanoid Robotics
Get the latest developments, breakthroughs, and insights in humanoid robotics — delivered straight to your inbox.




