Cross-Embodiment Robot Foundation World Models with Latent Actions
What happened
arXiv:2610.10846v1 Announce Type: new Abstract: The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. This unified action space improves the world model's performance when adapted to previously unseen robot embodiments.
LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM's downstream performance to scale positively with the number of embodiments used during pretraining.
In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlight the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics.
Sources & evidence
- arXiv Robotics (cs.RO) Reporting source
Cross-Embodiment Robot Foundation World Models with Latent Actions ↗
https://arxiv.org/abs/2610.10846