Published on

Robotera’s VPP2 Tops RoboDojo Simulation Benchmark by Linking Video Prediction to Robot Actions

Get your news fromHumanoids Daily

One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

P.A.
Written byP.A.
  • VPP2 leads RoboDojo-Sim with an average score of 39.26 and a 32.26% success rate; RoboDojo’s real-world leaderboard is separate.
  • The model combines video prediction and action generation, using a visual planner to help translate instructions into robot movements.
  • Robotera separately reports 58.5% average success across ten ALOHA task categories. Public code and checkpoints allow researchers to explore the approach.

Robotera’s Video Prediction Policy 2, or VPP2, has taken first place on RoboDojo’s simulation leaderboard, using a model that connects predictions of how a scene will change with the actions a robot should take.

The result, added to the benchmark on October 8 and announced by Robotera the following day, puts VPP2 ahead of PhysicalRSI on both average score and task success rate. The team has also published code, model checkpoints and evaluation instructions.

RoboDojo leaderboard illustration showing VPP2 and other robot policies on a mountain, dated October 11, 2026.
RoboDojo’s leaderboard illustration, dated October 11, 2026. Image: RoboDojo.

What the leaderboard measures

The official RoboDojo benchmark evaluates robot manipulation across capabilities including generalization, precision, long-horizon tasks, memory and open-ended tasks. Its simulation and real-world results are reported separately.

As checked on October 11, VPP2’s simulation score of 39.26 compares with 36.27 for second-placed PhysicalRSI. Their success rates are closer: 32.26% versus 31.38%, a difference of 0.88 percentage points. Score and success rate are distinct measures, and neither should be confused with reliability in a customer deployment.

The weekly humanoid robotics briefing

One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.

Read recent issues

The detailed results also show that leadership overall does not mean leadership on every capability. VPP2’s open-ended-task success rate is 4.08%, compared with PhysicalRSI’s 24.92%. That uneven performance is useful context for what remains difficult even for the leading systems.

Predicting a scene, then choosing an action

VPP2 is a world-action model: it connects a prediction of future visual states with a policy for acting. According to the research project, the pipeline starts by further training a video foundation model on manipulation footage with detailed captions.

The researchers then post-train and distill that model into a visual planner that makes its prediction in a single step. An action module learns to translate the visual information into robot actions. The aim is to make video prediction useful for executing an instruction, rather than only generating plausible-looking footage.

Robotera says the leaderboard result was achieved without additional data or agent-based self-improvement. That claim should not be read as an absence of training: the published method explicitly includes video pretraining, post-training and action learning.

Separate tests on physical hardware

Beyond RoboDojo, Robotera reports an average success rate of 58.5% across ten task categories on the ALOHA platform, and 45.0% on LIBERO-Pro. It also reports that adding high-level planning increased average success across five task groups from 27.6% to 57.6%. These are separate evaluations, rather than evidence that VPP2 leads RoboDojo’s real-world track.

The official repository provides RoboDojo and LIBERO training and evaluation guides, along with links to video and action checkpoints. The team announced the release of its Stage-1 and Stage-2 video-model weights on October 10. Code is MIT-licensed; model weights and other assets have their own terms.

For researchers, that makes the release more useful than a leaderboard announcement alone: there are implementation materials to examine and a concrete approach to test across other manipulation tasks.

Share this article

The weekly humanoid robotics briefing

One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.

Read recent issues