Fiatlux: A Long-Horizon Benchmark for Humanoid Ladder Climbing and Light-Bulb Replacement
What happened
arXiv:2609.38216v1 Announce Type: new Abstract: Existing benchmarks evaluate tabletop manipulation, flat-floor household activity, or humanoid locomotion and manipulation as separate task groups; none scores vertical mobility and dexterous work on a fragile payload in one long-horizon episode. Runs that fall short can earn partial credit.
The benchmark code and the teleoperated recordings are available at fiatlux-bench.github.io. In one episode, a Unitree G1 humanoid positions a step ladder under a ceiling or wall fixture, climbs it, exchanges a spent bulb in a socket for a fresh one, and leaves the spent one in a disposal crate.
The goal is a successful replacement, with the fresh bulb seated, the spent one disposed of, neither dropped, and a fragility bound not crossed. Observations are split into a standard mode (signals a physical robot could sense or estimate) and a privileged mode (exact simulator state).
Sources & evidence
- arXiv Robotics (cs.RO) Reporting source
Fiatlux: A Long-Horizon Benchmark for Humanoid Ladder Climbing and Light-Bulb Replacement ↗
https://arxiv.org/abs/2609.38216