Published on

Rhoda Tests Whether Scaling Web-Video Pretraining Makes Robots Better at Real Work

Get your news fromHumanoids Daily

One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

Humanoids Daily
Written byHumanoids Daily

Supported by RoboStrategy

Sponsored

Investing involves risk, including possible loss of principal. Read the prospectus before investing.

Rhoda’s two-arm robot at a bearing-unpacking workstation, with an inset showing bearings in a tote.
Rhoda’s bearing-unpacking demonstration. Image: Rhoda AI.
  • Rhoda is testing the link between video pretraining and useful robot work.
  • Its Direct Video-Action system predicts future video, then translates those predictions into movement.
  • Video models, sensorized demonstrations and simulation address different parts of the robotics learning problem—and can be combined.

Rhoda AI reports improved industrial manipulation from scaled web-video pretraining. Its September 10 study, shared on X today, tests a central physical-AI premise.

The result: completion with a deadline

Across four model sizes, Rhoda reports 3.7%, 65.0%, 75.3% and 84.7% completion without intervention in under 100 seconds, unpacking bearings and sorting waste. The largest scored 94/111. Larger models also used more compute and video, so this comparison does not isolate parameter count.

A separate fixed-size experiment increased pretraining compute: performance rose from 57.8% to 75.3%, plateauing near the upper budgets with full task data. Differences widened with reduced demonstrations. Better held-out video prediction also correlated with robot performance.

The company’s 200-plus evaluation hours cover one task and setup, with one post-training run per condition. Its deadline is stricter than the customer requires.

The weekly humanoid robotics briefing

One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.

Read recent issues

How video becomes a robot policy

Rhoda’s approach builds on the Direct Video-Action, or DVA, architecture we covered at its March launch.

The video model predicts what the robot should see next, using its observation history and other inputs. A separate inverse dynamics model works backward from that visual future to the movements needed to produce it. The robot acts, observes the result and repeats the cycle.

Rhoda pretrains its causal video model from scratch on general web video, then adapts it using robot task demonstrations. “Causal” means predictions follow from preceding observations. Its March technical account describes overlapping prediction and execution so the robot can keep moving while the next prediction is computed.

This gives video a direct role in choosing behavior. It also leaves an essential translation problem: predicting a useful future and making a particular machine reach it require different capabilities.

Where the other paths diverge

Our recent UMI feature examined another response to the data problem: designing what people wear or hold while demonstrating skills. Matching the collection device to the robot can make those demonstrations easier to transfer, while recording movement and sensing information that an ordinary video may lack.

These approaches can be complementary. Broad pretraining can supply reusable representations; purpose-built demonstrations can supply precise examples of how to act. The distinction is where a company invests to make the next useful skill cheaper to acquire.

Some of the companies we have covered recently illustrate several combinations:

CompanyEmphasis in its published approach
Odyssey / FlexionA broadly pretrained world model, adapted into humanoid policies through Flexion’s robot-learning and control work.
MimicA pretrained Cosmos video backbone; partially denoised visual representations guide an action decoder. It is a close comparison for video-based control, with a different implementation.
Generalist AIA foundation model built around large-scale physical interaction data, with approximately 99% of GEN-1’s parameters trained from scratch.
Genesis AIModel, dexterous hand, tactile collection glove and simulation developed together, linking the learning strategy to hardware.
Physical IntelligenceA VLA approach that incorporates diverse data and prompts. Its π0.7 work includes language coaching and world-model-generated visual subgoals.
Skild AISimulation and internet video used to pretrain a model intended to transfer across different robot bodies.

The table describes emphases, not exclusive categories or a ranking. Physical Intelligence’s use of visual subgoals, for example, shows why “VLA versus world model” can obscure how systems are actually assembled. Likewise, training from scratch says where parameters begin, not what kind of data a model learns from.

For readers tracking this field, the useful comparison is increasingly concrete: what information enters pretraining, what additional demonstrations each task needs, how predictions become controls, and how reliably the resulting robot works. A shared ambition for general-purpose robotics leaves plenty of room for different answers to each question.

Sources and further reading

Share this article

The weekly humanoid robotics briefing

One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.

Read recent issues