WORLDTECH NEWS Global technology intelligence.Contact
← Back to WORLDTECH
Robotics SINGLE SOURCE

SUAVE: Unified Video-Action Models via Masked Diffusion

A robotic hand reaching into a digital network on a blue background, symbolizing AI technology.
Illustrative photo.Photo by Tara Winstead on Pexels

What happened

arXiv:2610.04009v1 Announce Type: new Abstract: Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting.

World action models (WAMs) built on video diffusion backbones can imagine but treat language as frozen conditioning on a continuous latent space. Unified models bring these modalities into one architecture, but they either decode autoregressively, one token at a time, or keep video continuous with an auxiliary action head.

Choosing which tokens (the small pieces of text a model reads and writes) to mask at inference (running a trained model to get an answer, rather than training it) turns the same network into a world model, a robot policy, or a video-action model. For action-free co-training, the action positions of unlabeled video are filled with mask tokens and excluded from the loss.

Sources & evidence