- Published on
Rhoda Tests Whether Scaling Web-Video Pretraining Makes Robots Better at Real Work
One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

Supported by RoboStrategySupport for Humanoids Daily comes from RoboStrategy
SponsoredInvesting involves risk, including possible loss of principal. Read the prospectus before investing.

- Rhoda is testing the link between video pretraining and useful robot work.
- Its Direct Video-Action system predicts future video, then translates those predictions into movement.
- Video models, sensorized demonstrations and simulation address different parts of the robotics learning problem—and can be combined.
Rhoda AI reports improved industrial manipulation from scaled web-video pretraining. Its September 10 study, shared on X today, tests a central physical-AI premise.
The result: completion with a deadline
Across four model sizes, Rhoda reports 3.7%, 65.0%, 75.3% and 84.7% completion without intervention in under 100 seconds, unpacking bearings and sorting waste. The largest scored 94/111. Larger models also used more compute and video, so this comparison does not isolate parameter count.
A separate fixed-size experiment increased pretraining compute: performance rose from 57.8% to 75.3%, plateauing near the upper budgets with full task data. Differences widened with reduced demonstrations. Better held-out video prediction also correlated with robot performance.
The company’s 200-plus evaluation hours cover one task and setup, with one post-training run per condition. Its deadline is stricter than the customer requires.
The weekly humanoid robotics briefing
One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.
Read recent issuesHow video becomes a robot policy
Rhoda’s approach builds on the Direct Video-Action, or DVA, architecture we covered at its March launch.
The video model predicts what the robot should see next, using its observation history and other inputs. A separate inverse dynamics model works backward from that visual future to the movements needed to produce it. The robot acts, observes the result and repeats the cycle.
Rhoda pretrains its causal video model from scratch on general web video, then adapts it using robot task demonstrations. “Causal” means predictions follow from preceding observations. Its March technical account describes overlapping prediction and execution so the robot can keep moving while the next prediction is computed.
This gives video a direct role in choosing behavior. It also leaves an essential translation problem: predicting a useful future and making a particular machine reach it require different capabilities.
Where the other paths diverge
Our recent UMI feature examined another response to the data problem: designing what people wear or hold while demonstrating skills. Matching the collection device to the robot can make those demonstrations easier to transfer, while recording movement and sensing information that an ordinary video may lack.
These approaches can be complementary. Broad pretraining can supply reusable representations; purpose-built demonstrations can supply precise examples of how to act. The distinction is where a company invests to make the next useful skill cheaper to acquire.
Some of the companies we have covered recently illustrate several combinations:
| Company | Emphasis in its published approach |
|---|---|
| Odyssey / Flexion | A broadly pretrained world model, adapted into humanoid policies through Flexion’s robot-learning and control work. |
| Mimic | A pretrained Cosmos video backbone; partially denoised visual representations guide an action decoder. It is a close comparison for video-based control, with a different implementation. |
| Generalist AI | A foundation model built around large-scale physical interaction data, with approximately 99% of GEN-1’s parameters trained from scratch. |
| Genesis AI | Model, dexterous hand, tactile collection glove and simulation developed together, linking the learning strategy to hardware. |
| Physical Intelligence | A VLA approach that incorporates diverse data and prompts. Its π0.7 work includes language coaching and world-model-generated visual subgoals. |
| Skild AI | Simulation and internet video used to pretrain a model intended to transfer across different robot bodies. |
The table describes emphases, not exclusive categories or a ranking. Physical Intelligence’s use of visual subgoals, for example, shows why “VLA versus world model” can obscure how systems are actually assembled. Likewise, training from scratch says where parameters begin, not what kind of data a model learns from.
For readers tracking this field, the useful comparison is increasingly concrete: what information enters pretraining, what additional demonstrations each task needs, how predictions become controls, and how reliably the resulting robot works. A shared ambition for general-purpose robotics leaves plenty of room for different answers to each question.
Sources and further reading
- Rhoda: Scaling Web-Video Pre-training
- Rhoda: Direct Video-Action technical explanation
- Rhoda’s X post
- Odyssey: Introducing Odyssey-3
- Mimic: mimic-video research
- Generalist: Going Beyond World Models & VLAs
- Genesis AI: GENE-26.5 announcement
- Physical Intelligence: A Steerable Model with Emergent Capabilities
- Skild: Building the General-Purpose Robotic Brain
Share this article
Read next
- Published on
- Reading time
- 5 min read
Rhoda AI Hits $1.7B Valuation, Unveils "Direct Video-Action" Model to Bridge the Real-World Gap
- Published on
- Reading time
- 4 min read
Rhoda AI Breaks Stealth: Jagdeep Singh Teases "Real World" Generalization
- Published on
- Reading time
- 3 min read
Stealth Startups Emerge With Over $300 Million to Join Crowded Humanoid Robot Field
- Published on
- Reading time
- 8 min read
Uncaged: Agility CEO Peggy Johnson and Co-Founder Jonathan Hurst Detail Digit 5 Architecture, AI Stack, and Supply Chains
- Published on
- Reading time
- 5 min read
Agility unveils Digit 5, designed to work closer to people
- Published on
- Reading time
- 7 min read
Reward AI Exits Stealth with OM-1, Pitching Zero-Shot Human-to-Robot Manipulation
The weekly humanoid robotics briefing
One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.
Read recent issues
















