Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation
What happened
arXiv:2610.02368v1 Announce Type: new Abstract: Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning.
Given the current observation and global instruction, the subgoal planner predicts a visual subgoal for the next subtask. The subgoal executor then jointly generates future visual trajectories and actions conditioned on the predicted subgoal.
Both components share a pretrained world-model representation, enabling task-level planning and action generation to benefit from common physical knowledge. On the RoboTwin Clean2Random benchmark, ViGAR achieves 82.00% and 67.02% success rates under the Clean and Random settings, respectively, surpassing the strongest baseline by 12.86 percentage points in average success rate.
Sources & evidence
- arXiv Robotics (cs.RO) Reporting source
Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation ↗
https://arxiv.org/abs/2610.02368