提出RTNav架构,让机器人在真实时间下高效导航找物
RTNav: Towards Real-Time Zero-Shot Object Navigation

- 将推理延迟、异步环境推进和计算约束纳入设计核心
- 在真实时间条件下,成功率达11%提升,任务效率提高5.1点
- 适合追求实时性能的智能体导航研究者
在未知环境中寻找未预见物体的导航任务,随着强大视觉与语言基础模型的发展已变得可行。然而,这些模型也带来了显著的推理延迟,在需要持续运行的现实世界中成为关键问题。当前多数先进方法基于同步仿真器开发,环境等待智能体行动,推理时间被视为免费。因此,智能体常按感知-推理-执行顺序设计,忽视时间限制。在真实时间执行中,墙钟时间计入任务预算,此类架构的低效性暴露无遗。我们发现,现有零样本物体导航方法在真实时间条件下性能显著下降。为此,提出RTNav:一种将推理延迟、异步环境推进和有限算力作为显式设计考量的简单但有效的架构。在真实时间版本的HM3D-v1、HM3D-v2和HM3D-OVON上评估,相比先前工作,成功率最高提升11%,成功加权完成时间最高提升5.1分。
原文摘要 · Abstract (English)
Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language foundation models. However, these models also introduce non-negligible inference latency, which becomes an important concern when agents must operate continuously in the real world. Most state-of-the-art methods are still developed in synchronous simulators, where the environment waits for the agent to act and inference time is effectively free. As a result, agents are often designed around the sequential execution of perception, reasoning, and action, with little regard for time constraints. Under real-time execution, where wall-clock time counts towards the task budget, the inefficiencies of these architectures become clear. We show that recent zero-shot object navigation methods suffer consistent performance degradation under such realistic timing conditions. Motivated by this observation, we propose RTNav, a simple but effective architecture that treats inference latency, asynchronous environment stepping, and bounded compute as explicit design considerations. Evaluated on real-time variants of HM3D-v1, HM3D-v2, and HM3D-OVON, RTNav improves the success rate by up to 11% and the Success weighted by Completion Time by up to 5.1 points over prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。