WORLDTECH NEWS Global technology intelligence.Contact
← Back to WORLDTECH
Robotics SINGLE SOURCE

SWAP: Stepwise Action Policy Routing for Vision-Language-Action Models

Aerial view of a drone flying over snow-covered mountains with dramatic clouds.
Illustrative photo.Photo by Alan Kabeš on Pexels

What happened

arXiv:2610.06926v1 Announce Type: new Abstract: Robot manipulation systems using Vision-Language-Action (VLA) model backbones typically use just one VLA for task execution. SWAP formulates policy routing as an offline reinforcement learning problem, learning a routing critic that selects the most appropriate policy at each decision step given the current observation.

However, individual VLAs do not perform well across different task states and environments. SWAP enables robots to select new policies to execute online rather than committing to a single policy for the duration of an episode. SWAP improves over fixed-policy execution and routing baselines, giving absolute improvements in real-world task success up to 33% while reducing successful trajectory robot action step length by 28.3%.

Sources & evidence