NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
What happened
arXiv:2610.10787v1 Announce Type: new Abstract: Language models trained with long-horizon agentic (AI that carries out multi-step tasks rather than answering one question) reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. With the complete harness, NavGPT-3 sets the state of the art on R2R-CE (81.51 SR) and, for the first time, brings an autonomous agent to human level: on RxR-CE it matches human followers in success (90.43 vs.
These results show that designing this embodied interface is central to connecting frontier language-model intelligence with low-level physical control. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency (the delay between asking for something and getting it) control. 77.7 nDTW) at 1 min 22 s per episode, versus roughly 3 min for a human.
Sources & evidence
- arXiv Robotics (cs.RO) Reporting source
NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime ↗
https://arxiv.org/abs/2610.10787