Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
What happened
arXiv:2610.10810v1 Announce Type: new Abstract: Long-horizon robotic manipulation is often built by chaining independently trained skills. These results suggest that some long-horizon composition failures are better addressed by restoring the scene before resuming the policy than by retrying from an off-support state.
Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. On a real Franka arm running a fine-tuned (further training of an existing model for a narrower job) $\pi_{0.5}$ policy, the same monitor is limited by exterior-camera observability, yet closing the loop still recovers some otherwise-terminal failures, motivating wrist and gripper sensing.
Sources & evidence
- arXiv Robotics (cs.RO) Reporting source
Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams โ
https://arxiv.org/abs/2610.10810