WORLDTECH NEWS Global technology intelligence.Contact
← Back to WORLDTECH
Robotics SINGLE SOURCE

DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

A white robotic arm operating indoors with a modern design and advanced technology.
Illustrative photo.Photo by Magda Ehlers on Pexels

What happened

arXiv:2610.00360v1 Announce Type: new Abstract: Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise that serves the first goal can interfere with the second. DexPolicy makes that scale an explicit function of training steps, annealing from broad to narrow exploration while holding loss, architecture, reward, and optimizer settings fixed.

Across five YCB objects and three training seeds, mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO). On a RealMan RM75 arm with an Inspire/RH56 hand, 360 trials over three objects raise mean Target success from 25.0% to 85.0% (FPO), 10.0% to 63.3% (GRPO), and 8.3% to 43.3% (PPO), with one trained model per object-method condition.

Sources & evidence