Published on

Unitree Open-Sources UnifoLM-WLA-1.0 to Tackle Humanoid Generalization

Get your news fromHumanoids Daily

One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

Humanoids Daily
Written byHumanoids Daily
A Unitree G1 humanoid robot walking toward an open front-loading washing machine while holding green laundry in its left hand, alongside picture-in-picture head and wrist camera feeds and real-time policy execution terminal logs showing sub-105 millisecond inference latency.
Whole-body mobile manipulation in action: A Unitree G1 autonomously carries laundry toward a washing machine during real-robot evaluations of UnifoLM-WLA-1.0. Inset feeds show stereo head and wrist vision alongside terminal telemetry logging sub-105 ms policy inference latencies. Image: Unitree Robotics
  • Unitree Robotics has fully open-sourced UnifoLM-WLA-1.0, a 6-billion-parameter Vision-Language-Action (VLA) foundation model designed for end-to-end humanoid manipulation.
  • The architecture pairs an embodied reasoning backbone (UnifoLM-ER-1-4B) with an MMDiT flow-based action expert trained on approximately 2,500 hours of real-world robot data, including the BitRobot-HIW-500 dataset.
  • The model unifies desktop manipulation and mobile whole-body coordination into a single policy checkpoint, evaluating across 64 real-robot tasks on the Unitree G1 platform.
  • Released amid mounting industry debates over whether frontier LLMs like GPT-6 Astra will commoditize physical AI layers, the open release reinforces the open-source embodied commons following NVIDIA’s recent $12.9 billion acquisition of Hugging Face.

Having scaled its bipedal manufacturing lines to more than 18,000 cumulative units and staged a blockbuster public debut in Shanghai, Unitree Robotics is targeting the software bottleneck that keeps most of those machines confined to labs.

The Hangzhou-based manufacturer has fully open-sourced UnifoLM-WLA-1.0, a 6-billion-parameter general-purpose humanoid foundation model designed to bridge the gap between multimodal spatial reasoning and physical execution. The release marks a major escalation in the race for open physical AI, introducing a unified architecture that coordinates tabletop dexterity and mobile, whole-body manipulation from a single model checkpoint.

The software drop follows hot on the heels of Unitree’s UnifoLM-X2 dynamic world model, which powered autonomous bipedal sparring demonstrations just days earlier. But while X2 prioritized low-latency, contact-rich locomotion, UnifoLM-WLA-1.0 focuses on the industrial and domestic tasks where humanoid manipulation has traditionally stumbled: grasping, tidying, folding, and tool handling across variable environments.

Inside the UnifoLM Architecture

Rather than treating robot control as a basic reinforcement learning policy or a narrow imitation learning loop, UnifoLM-WLA-1.0 is built around a structured three-stage pipeline that marries visual understanding with dynamic physical forecasting.

Stay ahead in humanoid robotics

One email a week: the launches, funding and research that mattered. Join the researchers, founders and technologists following the field.

At its foundation sits UnifoLM-ER-1-4B, an embodied reasoner based on Qwen3-VL-4B. Co-trained on standard image-text pairs and more than 5 million embodied samples—spanning 3D bounding boxes, 2D trajectory prediction, and multi-image spatial question-answering—the model is engineered to preserve broad multimodal understanding while sharpening fine-grained physical awareness. In internal benchmark evaluations released by Unitree, the 4B parameter ER model outperformed several larger open-source baselines, including RoboBrain2.0-7B, Pelican-7B, and Cosmos-R1-7B, across tests such as Spatial-Bench, Pixmo-Point, and BLINK.

+-------------------------------------------------------------------+
|                        UnifoLM-ER-Flow                            |
|  (Multimodal Reasoning + VQ-VAE Dynamic Mask + Discretized RVQ)   |
+---------------------------------+---------------------------------+
|
v
+-----------------------------------+
|        Action Expert MMDiT        |
|           (Flow Decoder)          |
+-----------------+-----------------+
|
+---------------------+---------------------+
|                                           |
v                                           v
[Discrete Action Tokens]                  [Continuous Actions]

End-Effector Trajectories               - 10 Whole-Body Tasks

Dexterous Hand / Gripper                - 54 Tabletop Tasks

Lower-Body Locomotion

Building on the reasoning core, Unitree introduces UnifoLM-ER-Flow, which integrates interaction-centric world modeling via dynamic region prediction. The system estimates optical flow between timesteps to capture scene changes induced by physical contact, using a Vector Quantized-Variational AutoEncoder (VQ-VAE) to convert these dynamic masks into discrete tokens. When given an initial camera frame and a task prompt, the model predicts how the physical scene should mutate over future steps.

Finally, action generation is divided across a unified action space partitioned into three distinct streams: end-effector (EEF) poses, hand or gripper joints, and lower-body joints. Each component is discretized using Residual Vector Quantization (RVQ). To generate continuous trajectories, the architecture deploys an MMDiT (Multi-Modal Diffusion Transformer) flow-matching decoder, enabling the network to output smooth, high-frequency motor commands conditioned on the ER-Flow representations.

Scaling Manipulation Beyond the Bench

Unitree trained UnifoLM-WLA-1.0 on roughly 2,500 hours of real-robot demonstration data, drawing from internal datasets as well as open community repositories like BitRobot-HIW-500—the massive household teleoperation dataset hosted on Hugging Face.

The primary target platform is the company’s sub-$30,000 Unitree G1 humanoid. In demonstration footage published on the project repository, a single model instance operates across 64 real-world manipulation tasks without requiring task-specific fine-tuning:

  • 10 Whole-Body Manipulation Tasks: Mobile tasks requiring coordinated bipedal locomotion, torso posture adjustment, and dual-arm reach, including taking out trash, putting clothes into a washing machine, making beds, and shelving inventory.
  • 54 Tabletop Manipulation Tasks: High-dexterity static operations such as folding towels, sorting small components from conveyor belts, uncapping containers, and plugging in power cords.

Crucially, the system supports both basic parallel grippers and two distinct types of multi-finger dexterous hands. Cross-end-effector generalization has historically been a significant barrier for robotics developers; swapping out a two-jaw clamp for a five-fingered hand typically breaks policy representations and requires comprehensive data recollection. By learning a shared latent representation across multiple embodiments and end-effectors, UnifoLM aims to make policy transfer far less fragile.

Open Commons in the Frontier Era

The release arrives amid an intensifying strategic debate across the embodied AI landscape. As demonstrated by recent experiments with GPT-6 Astra learning to paint in real life, frontier generalist models are increasingly proving capable of zero-shot physical planning and tool-use, leading some researchers to wonder whether specialized physical AI architectures will ultimately be swallowed by general reasoning APIs.

Yet hardware advocates argue that algorithms commoditize rapidly, leaving competitive advantages in the hands of mass-production platforms and open data flywheels. Unitree's decision to open-source UnifoLM-WLA-1.0 fits into a broader expansion of open robotics infrastructure—highlighted by NVIDIA’s $12.9 billion buyout of Hugging Face to solidify open model hosting and LeRobot community standards.

The release also addresses the sobering operational realities voiced by Unitree founder and CEO Wang Xingxing. Speaking at the World Robot Conference, Wang gave a frank assessment of why humanoids still aren't ready to scale in production environments, pointing directly to the "lossy" nature of physical contact. Unlike language models operating in lossless vector spaces, physical robots suffer from compounding millimeter-scale deviations that collapse success rates when novel environments deviate from training distributions.

Wang argued that the industry's genuine "ChatGPT moment" won't arrive until a humanoid can enter an unfamiliar room and successfully complete 80% of voice-prompted tasks—a milestone he projected was still two to ten years away.

UnifoLM-WLA-1.0 is Unitree's attempt to accelerate that timeline through public collaboration. By open-sourcing the weights and training recipes, Unitree is extending its familiar hardware playbook to the algorithmic layer: just as it undercut Western hardware pricing by commercializing low-cost actuators, it is now lowering the barrier to deployable, full-body foundation policies. Whether 2,500 hours of multi-source data and an MMDiT flow decoder can reliably solve the "final millimeter" in unconstrained environments remains an open question, but Unitree has ensured the entire research community can now test the hypothesis.

Share this article

Stay ahead in humanoid robotics

One email a week: the launches, funding and research that mattered. Join the researchers, founders and technologists following the field.