- Published on
Unitree’s AI models explained: UniFoLM, WLA and X2
One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

Supported by RoboStrategySupport for Humanoids Daily comes from RoboStrategy
SponsoredInvesting involves risk, including possible loss of principal. Read the prospectus before investing.

A robot throwing a punch and a robot putting clothes into a washing machine might seem to have little in common. For Unitree, they are two ways of approaching the same problem: how to make a machine respond to the world around it, rather than depend on a person directing every movement.
The sparring clip attracts the attention. The laundry task points toward the business opportunity. Selling humanoid bodies is one thing; making them useful in homes and workplaces requires software that can interpret instructions, handle unfamiliar objects and recover when an action does not go as planned.
Unitree Robotics is pursuing that software through UniFoLM, a family of robot AI models. Recent announcements have introduced UnifoLM-X2-1.0 and UnifoLM-WLA-1.0, joining earlier projects with similarly cryptic names. They share a research direction, but they do different jobs—and their public availability varies considerably.
For readers trying to understand what is actually new, those distinctions matter more than the version numbers.
The weekly humanoid robotics briefing
One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.
Read recent issuesWhat is UniFoLM?
UniFoLM is the umbrella name Unitree uses for a family of robot foundation-model projects. It is not a single app, nor does the name identify the software behind every Unitree demonstration.
A foundation model learns patterns from a broad body of training data that can support multiple tasks. In robotics, those patterns must eventually help a physical machine act: grasp an object, move an arm, or coordinate its body with what its cameras see.
Unitree’s public projects explore different parts of that problem. This is a map of the main models covered here, rather than a claim that each directly replaces the one before it:
| Model | Main focus | What to associate it with |
|---|---|---|
| UnifoLM-WMA-0 | Predicting physical interactions and using those predictions for robot learning | World-model simulation and action selection |
| UnifoLM-VLA-0 | Connecting visual observations and instructions to actions | Manipulation tasks on the G1 |
| UnifoLM-WLA-1.0 | Combining understanding, interaction prediction and body control | Tabletop and mobile whole-body manipulation |
| UnifoLM-X2-1.0 | Fast decisions during dynamic interaction | Unitree’s autonomous sparring demonstration |
The first two are documented in Unitree’s WMA repository and VLA repository. WLA has its own project page; X2 was introduced through a company demonstration.
Predicting what happens next: WMA
A world model tries to predict how a scene will change. Imagine a robot reaching for a cup. Recognizing the cup is useful, but predicting what happens when the gripper pushes against its edge is a different ability.
Unitree’s UnifoLM-WMA-0, short for World-Model-Action, uses prediction in two ways. It can generate simulated interactions for training, and it can help an action-producing system choose what to do next. The official release includes experiments involving both Unitree’s Z1 arms and the G1 humanoid.
Unitree dates the release of its training code, inference code and model weights to September 15, 2025, with deployment code following a week later. That makes WMA an established public research project within the family, rather than a new name introduced alongside September 2026’s demonstrations. Official WMA documentation.
Turning instructions into movements: VLA
A vision-language-action model, or VLA, connects what a robot sees and what it is asked to do with physical actions.
Consider the instruction “put the pencil in the box.” A useful system must identify both objects, understand their positions and produce movements that complete the task. A written answer describing the correct procedure is not enough.
UnifoLM-VLA-0 is Unitree’s earlier manipulation-focused model. Its documentation describes a single policy tested across 12 categories of tasks, with datasets covering activities such as cleaning a table, folding a towel and organizing tools. Unitree released training and inference code and model weights on January 29, 2026. Official VLA documentation.
World models and VLAs are not necessarily competing ideas. One describes an ability to predict; the other describes a route from perception and instructions to action. A robot system can bring both together.
Bringing the body along: WLA
That combination is central to UnifoLM-WLA-1.0. Unitree describes a six-billion-parameter model trained on roughly 2,500 hours of real-robot data, with evaluations spanning 54 tabletop tasks and 10 whole-body tasks.
The wider ambition is to connect manipulation with movement through a room. A robot working at a table can keep much of its body in place. Handling laundry or putting something on a shelf requires it to coordinate reaching, posture and locomotion as well.
WLA brings together three components: ER-1 for embodied reasoning, ER-Flow for representations that include predicted changes and actions, and an action expert that generates movements. ER-1 and ER-Flow are components of this system, not separate humanoid products. Unitree’s technical overview.
The reported task count describes Unitree’s evaluations. It does not establish that the robot can reliably perform arbitrary chores in an unfamiliar home. For the architecture and release details, read our coverage of UnifoLM-WLA-1.0.
Why show a robot sparring? X2
With UnifoLM-X2-1.0, Unitree chose a much faster interaction to showcase. The company describes its footage of a G1 sparring with a padded human partner as autonomous and driven by a real-time world model.
The research question is easy to see: can the system make decisions quickly enough while both the robot and the person are moving? That is a different demonstration from placing an object on a table, even if both involve predicting what will happen next.
The distinction between autonomy and teleoperation is also central. In teleoperation, a person directs the robot’s movements. Unitree presents X2 as selecting its actions itself during this demonstration.
That is the company’s account of the system, not an independent test of its reliability. Sparring footage does not establish safe operation around unprotected bystanders, or prove that the same capability transfers to factory work. Our X2 report examines the demonstration in more detail.
Which models can you actually download?
Availability checked September 16, 2026. A model announcement and a complete public release are different milestones.
| Project | Public release status |
|---|---|
| WMA-0 | Training, inference and deployment code, with linked model checkpoints in the official repository. |
| VLA-0 | Training and inference code, with linked checkpoints in the official repository. |
| WLA-1.0 | ER-1 and ER-Flow weights released. WLA-Base and post-training code remain pending on the release checklist. |
| X2-1.0 | Demonstrated publicly; we have not verified an official code-and-weights release. |
For WLA, downloading the released reasoning components does not give a developer the complete action policy shown in the videos. For any model, reproducing a demonstration also requires compatible hardware, sensors, software and deployment configuration.
Nor should a buyer assume that purchasing a G1 includes all these capabilities. Research releases describe particular systems and experiments; they are not a specification of the software supplied with every retail robot.
The question beyond the demonstrations
The range of UniFoLM projects shows why “robot intelligence” is difficult to reduce to one score. Understanding where an object is, predicting how it will move and controlling a body around it are connected challenges. Progress on one does not automatically solve the others.
For Unitree, there is a commercial question behind that research: can the software make its hardware useful enough for customers to put it to work? The company’s own CEO has discussed the obstacles that still keep humanoids out of routine factory work.
The evidence to watch is what happens beyond a prepared demonstration: whether a robot completes the task repeatedly, adapts to a changed setting and needs less intervention from its operators. A sparring clip can make people notice the AI. Dependable everyday work is what would make them rely on it.
Share this article
Read next
- Published on
- Reading time
- 3 min read
Unitree launches G1+ with stronger motors, a moving neck and a $15,000 price
- Published on
- Reading time
- 7 min read
Unitree Begins Open Release of UnifoLM-WLA-1.0 to Tackle Humanoid Generalization
- Published on
- Reading time
- 5 min read
Unitree Unveils UnifoLM-X2: World Model AI Powers Fully Autonomous Robot Combat
- Published on
- Reading time
- 6 min read
Hack One Robot, Reach the Next: How Unitree’s G1 Left the Door Open to Root Takeover
- Published on
- Reading time
- 7 min read
Not Ready to Scale: Unitree CEO Wang Xingxing on Why Humanoids Still Aren’t Working in Factories
- Published on
- Reading time
- 6 min read
Global Humanoid Shipments Surpass 22,000 in H1 2026 as AGIBOT Leads and Form-Factor Debates Intensify
The weekly humanoid robotics briefing
One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.
Read recent issues













