arXiv:2508.04598cs.RO2025-08被引 26

让机器人理解复杂指令,在真实环境里找东西并导航到目标。

$NavA^3$: Understanding Any Instruction, Navigating Anywhere, Finding Anything

  • 分两阶段框架:先理解指令定位大区域,再用点云模型精确定位物体。
  • 在真实场景中完成长时程导航任务,性能超越现有方法。
  • 适合研究通用机器人导航、具身智能的学者和开发者。

具身导航是机器人在物理环境中移动与交互的基础能力。然而,现有导航任务多聚焦于预定义目标或简单指令跟随,与真实世界中复杂的开放式场景需求存在显著差距。为此,我们提出一项具有挑战性的长时程导航任务,要求理解高层级人类指令,并在真实环境中进行空间感知的开放词汇目标导航。现有方法因难以理解高层级指令及在开放词汇下定位物体而表现受限。本文提出 $NavA^3$,一个分两阶段的层次化框架:全局策略利用 Reasoning-VLM 解析高层指令,并结合全局 3D 场景视图进行推理,引导至最可能包含目标物体的区域;局部策略基于自建的 100 万样本空间感知物体属性数据集训练 NaviAfford 模型(PointingVLM),实现鲁棒的开放词汇目标定位与空间感知,以精准识别目标并导航。大量实验表明,$NavA^3$ 在导航性能上达到当前最优,在不同机器人形态的真实场景中均能成功完成长时程导航任务,为通用具身导航铺平道路。数据集与代码将公开。项目网站:https://NavigationA3.github.io/。

原文摘要 · Abstract (English)

Embodied navigation is a fundamental capability of embodied intelligence, enabling robots to move and interact within physical environments. However, existing navigation tasks primarily focus on predefined object navigation or instruction following, which significantly differs from human needs in real-world scenarios involving complex, open-ended scenes. To bridge this gap, we introduce a challenging long-horizon navigation task that requires understanding high-level human instructions and performing spatial-aware object navigation in real-world environments. Existing embodied navigation methods struggle with such tasks due to their limitations in comprehending high-level human instructions and localizing objects with an open vocabulary. In this paper, we propose $NavA^3$, a hierarchical framework divided into two stages: global and local policies. In the global policy, we leverage the reasoning capabilities of Reasoning-VLM to parse high-level human instructions and integrate them with global 3D scene views. This allows us to reason and navigate to regions most likely to contain the goal object. In the local policy, we have collected a dataset of 1.0 million samples of spatial-aware object affordances to train the NaviAfford model (PointingVLM), which provides robust open-vocabulary object localization and spatial awareness for precise goal identification and navigation in complex environments. Extensive experiments demonstrate that $NavA^3$ achieves SOTA results in navigation performance and can successfully complete longhorizon navigation tasks across different robot embodiments in real-world settings, paving the way for universal embodied navigation. The dataset and code will be made available. Project website: https://NavigationA3.github.io/.

具身智能导航指令理解开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。