DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
What happened
arXiv:2610.00317v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation.
While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies.
Key facts
- Sequence-level reinforcement learning can address this limitation, but typically — requires: policy rollouts and closed-loop interaction, which are costly for real-robot manipulation
Sources & evidence
- arXiv Robotics (cs.RO) Reporting source
DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies ↗
https://arxiv.org/abs/2610.00317