Motus2 is here: a self-evolving general world model for dexterous robot manipulation.
A world model should do more than predict the next video frame. It should simulate the consequences of actions, judge which choice is better, and use those evaluations to keep improving its policy.
Getting a robot hand to move is not the same as giving it truly dexterous hands. A capable robot must learn from human manipulation, sense changes while interacting with objects, adjust its movements, and gradually build up practical skill.
Building on Motus, Motus2 learns from large-scale first-person human manipulation data. It progresses from monocular video to stereo video and human action, then transfers that experience to dexterous robot hands through robot trajectories and human-robot aligned data.
Motus2 brings action, prediction, and evaluation into a single shared-parameter model. It can generate candidate actions, simulate the outcomes they may produce, and determine which outcome is most useful for advancing the task. The resulting evaluations support both action selection during planning and policy updates during further training, forming a closed loop: generate actions → predict the future → evaluate outcomes → update the policy.
To handle contact in dexterous manipulation, Motus2 also introduces a lightweight tactile expert. It refines movements from the latest tactile feedback, giving the robot contact information beyond vision alone.
Imagine a robot learning to fry an egg. Simply copying the hand motion in a tutorial will fail when the egg slips or the pan moves. A thoughtful apprentice compares alternatives: Is this grip stable? What happens at a different angle? Which approach is more likely to work? During execution, it also adjusts force from finger contact; during continued learning, it turns judgments about different approaches into better skill.
Think through this step, then learn the next step better. The self-evolution pursued by Motus2 means continually using judgments about action outcomes to improve manipulation.
| Blog: https://www.shengshu.com/zh/motus-2/
| Project page: https://motus-robotics.github.io/motus2/
| Paper: https://arxiv.org/abs/2608.30237
Motus2 in three keywords
First: the first fully multimodal unified world model
Beyond vision, language, and action, Motus2 adds a tactile expert. The robot can not only see objects but also sense changing contact states for finer dexterous manipulation.
Second: the first closed-loop self-evolving iteration
One model serves as policy, simulator, and evaluator, allowing the robot to generate feedback with its own world model and use that feedback to improve its action policy.
Third: the first world model for high-DoF dexterous hands
Motus2 is pretrained on large-scale, human-only egocentric data, followed by over one hundred hours of robot-domain intermediate training and target-task fine-tuning. It has been deployed on tactile dexterous hands including the 22-DoF Sharpa Wave and the 20-DoF WUJI Hand 2.
A multi-layer Ego data system built from about 130,000 hours of first-person human data connects the way people manipulate objects to the way robots perform actions, transferring large-scale human experience to robots.
Real robots first: screwing in bulbs, tearing paper, turning pages, and finding objects
Before discussing the model, consider what the robot can do.
It aligns a bulb with its socket and screws it in until the light turns on; coordinates both hands to tear a paper towel; and remembers earlier information to find a square hidden beneath one of several cups.
These actions feel natural to people but are difficult for robots. With dexterous hands, the question is no longer merely whether an object can be picked up, but how to pinch and rotate it, when to release it, how much force to apply, and whether the contact state has changed.
Motus2 demonstrates page turning, can opening, dough rolling, multi-finger grasping, bulb installation, bimanual paper tearing, hidden-object search, and more. Some tasks demand two-hand coordination, others rely on continuing contact feedback, and some require remembering what was seen earlier.
View the official demo collection:
https://motus-robotics.github.io/motus2/#demo
One model closes the loop across action, prediction, and evaluation
Traditional imitation-learning robot policies mainly answer one question: given the current state, what should the robot do next?
When object position and orientation keep changing, the robot must also ask: what will happen if I take this action?
Motus2 answers through three capabilities in the same video-action model:
- Action (Policy / WAM): generates robot actions.
- Prediction (Simulator / AC-WM): predicts what happens after an action.
- Evaluation (Evaluator / VM): judges whether the result is closer to the task goal.
In short, it answers three questions at once: What should happen next? What will happen after the action? Is that outcome good? The robot begins to understand action consequences and judge action value instead of merely executing commands.
Training preserves a strict order: an action is generated only from information already available and cannot peek at the future; prediction then reads the candidate action; evaluation finally judges the action together with its predicted outcome. Normal control may generate actions alone, while planning or policy learning also invokes prediction and evaluation.
Putting all three capabilities in one model is only the first step. The important part is making prediction and evaluation improve action: during training, the model learns to propose better actions; during execution, it selects better candidates. Choosing correctly once and becoming more likely to choose correctly next time are two different forms of progress.
During training, the model generates candidate actions, predicts their future states, and uses the value model to determine which outcomes best support task completion. The scores then update the robot policy.
High-value actions become more likely and low-value actions less likely. Only action-related parameters are updated at this stage; prediction and evaluation remain fixed so the basis for judgment does not shift while the policy improves.
In other words, the future predicted by the model itself becomes feedback for training the model.
This also makes failure data more valuable.
If excessive grip force deforms a cup, the action should not be imitated—but the knowledge of what that action causes is exactly the experience a robot needs.
Successful trajectories tell the robot what to do. Failed and suboptimal trajectories help it understand what it did, what happened in the world, and why the task was not completed. The failed action is not worth copying, but the information it produced is worth keeping.
Before taking an action, Motus2 can also think one step further through Best-of-N: it proposes several candidate actions, predicts and scores their outcomes, and executes the better option. After an action segment, it reads the actual change in the scene and plans again. This inference-time reasoning changes the current choice but does not update model parameters.
On two real-robot tasks—phone placement and multi-finger manipulation—model-based reinforcement learning raised average success from 65% to 72.5%. Adding inference-time planning increased it to 75%, a 10-percentage-point gain over the base model.
130,000 hours of human data keep improving the model
Robot data are expensive and cannot scale as quickly as Internet video. Yet people interact with the physical world every day: picking up cups, opening caps, turning pages, and arranging objects all produce natural large-scale manipulation data.
Can a model first learn from these experiences, then transfer what it learns to a robot?
Motus2 builds a multi-layer Ego data system whose foundation is about 130,000 hours of first-person human data from monocular and stereo sources. First-person video shows, from the operator's own viewpoint, how hands interact with objects. Training proceeds in three stages:
First, observe how people act. Large-scale monocular Ego video teaches broad patterns of object change during grasping, movement, and contact, building basic world knowledge and manipulation priors.
Second, understand the hand-object relationship. Stereo video and action data add depth cues and finer-grained spatial and interaction information.
Finally, transfer the experience to a robot body. Human and dexterous robot hands differ in appearance and morphology. More than 100 hours of robot trajectories and human-robot aligned data adapt human experience to the robot's control space.
Experiments show that within the tested range, human-action prediction error keeps falling as stereo data increase from 2,000 to 20,000 hours. More high-quality human interaction data continue to improve the model.
Dexterity requires accurate vision, memory, and touch
Two capabilities are unavoidable in the complex physical world: memory and touch.
Suppose a square is hidden beneath one of three cups. Once the cups are rearranged, the current frame alone no longer reveals the target.
The robot must remember what happened earlier.
Motus2 therefore studies long-horizon information mechanisms. On tasks such as finding a hidden square and acting from historical cues, retaining the full history reaches a 57.5% average success rate. A hybrid approach that keeps recent and selected key observations while compressing the rest reaches 25.0%. More complete memory, however, also requires more computation and cache.
Vision alone is also insufficient for tasks such as tearing paper or extracting paper cups. The robot must know: Did contact happen? Is the grip secure? Did the object slip? Seeing accurately does not mean feeling accurately.
Motus2 therefore adds a lightweight tactile expert. The main model plans the action, while the tactile module reads the latest contact feedback before every short action segment and refines the motion. It reuses the main model's intermediate computation instead of rerunning the full video backbone. On paper-cup extraction and paper tearing, tactile feedback raises average success from 60% to 72.5%, an improvement of 12.5 percentage points.
84% average success across five real-robot tasks
A model's value must ultimately be tested on real robots. Across the five principal real-robot tasks in the paper, Motus2 achieves an average success rate of 84%.
One controlled comparison is especially notable. With the same target-task fine-tuning, a model pretrained only on human first-person data averages 51%. Adding robot-domain intermediate training raises success to 84%—a gain of 33 percentage points.
This means experience learned from human video can form a foundation for robot capability, while adaptation to the robot's own body and controls is what further improves real-world success. Seeing how people do it begins to become doing it with the robot's own hands.
The experiments were conducted on three bimanual robot configurations using WuJi Hand 1, WuJi Hand 2, and Sharpa Wave dexterous hands. Every capability ultimately returns to the physical world.
From acting in the world to beginning self-evolution
ShengShu Technology recently proposed a five-level roadmap for General World Models (GWM): L1 World Generation, L2 Interactive World, L3 Actionable World, L4 Autonomous World Agent, and L5 World Orchestrator.
These are not five independent model categories but a path of accumulating capability. From generating and understanding the world to acting in it, a robot must first solve how to act. The next step is to evaluate its own action outcomes and use real feedback to continually improve its policy.
Motus2 is a concrete embodied-intelligence implementation of this path. It already has the core L3 capability to understand the current state, predict the future after an action, and generate real robot movements. On top of that, Motus2 explicitly brings Evaluation and Feedback into a loop: action → prediction → evaluation → feedback → further learning.
The mechanism shows early characteristics of RSI (Recursive Self-Improvement): the model uses its own predictions and evaluations of the world to improve its action policy. The world model begins to demonstrate a degree of self-evolution, an important exploration from L3 toward the L4 autonomous world agent.
There is still a long way to go from multi-task policy optimization to long-term autonomous learning and continual evolution in an open world. Motus2 nevertheless validates an important step: as the system moves from L3's ability to act toward L4's ability to evolve autonomously, evaluation and feedback begin to enter the world-model loop.
Motus2 is an early attempt along that path.