SWAP: Stepwise Action Policy Routing for Vision-Language-Action Models
What happened
arXiv:2610.06926v1 Announce Type: new Abstract: Robot manipulation systems using Vision-Language-Action (VLA) model backbones typically use just one VLA for task execution. SWAP formulates policy routing as an offline reinforcement learning problem, learning a routing critic that selects the most appropriate policy at each decision step given the current observation.
However, individual VLAs do not perform well across different task states and environments. SWAP enables robots to select new policies to execute online rather than committing to a single policy for the duration of an episode. SWAP improves over fixed-policy execution and routing baselines, giving absolute improvements in real-world task success up to 33% while reducing successful trajectory robot action step length by 28.3%.
Sources & evidence
- arXiv Robotics (cs.RO) Reporting source
SWAP: Stepwise Action Policy Routing for Vision-Language-Action Models ↗
https://arxiv.org/abs/2610.06926