WORLDTECH NEWS Global technology intelligence.Contact
← Back to WORLDTECH
Robotics SINGLE SOURCE

NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime

A sleek white drone flying under a clear blue sky, capturing aerial views.
Illustrative photo.Photo by Bert Christiaens on Pexels

What happened

arXiv:2610.10787v1 Announce Type: new Abstract: Language models trained with long-horizon agentic (AI that carries out multi-step tasks rather than answering one question) reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. With the complete harness, NavGPT-3 sets the state of the art on R2R-CE (81.51 SR) and, for the first time, brings an autonomous agent to human level: on RxR-CE it matches human followers in success (90.43 vs.

These results show that designing this embodied interface is central to connecting frontier language-model intelligence with low-level physical control. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency (the delay between asking for something and getting it) control. 77.7 nDTW) at 1 min 22 s per episode, versus roughly 3 min for a human.

Sources & evidence