arXiv:2608.12707cs.RO2026-08

让机器人在陌生环境里听懂复杂指令,边走边看边判断。

SAP-Nav: Spatial Semantic Representation Meets Active Perception for Hierarchical Open-Vocabulary Object Navigation

论文配图:SAP-Nav: Spatial Semantic Representation Meets Active Perception for Hierarchical Open-Vocabulary Object Navigation
图 1 · 摘自论文原文
  • 通过主动观察构建可查询的空间语义地图,实时支持多层级目标定位。
  • 在区域级导航上比训练方法提升12.2%成功率,无需预先建图或训练。
  • 适用于真实机器人部署,支持从房间到实例的任意层级指令理解。

层次化开放词汇物体导航(OVON)要求智能体根据自由形式指令,在未见过的环境中定位目标,指令可能涉及场景、房间、区域或实例级别的提示。尽管最近工作LangMap已定义该任务,但在部分观测条件下可靠求解仍具挑战:空间定位需要持续的环境级证据,而目标验证则需清晰且具有区分性的候选视角。本文提出SAP-Nav,一个完全在线、零样本的框架,通过主动感知同时满足两项需求。SAP-Nav从主动获取的房间视图中逐步构建可查询的空间语义表示,使智能体可在任何已探索位置进行空间语义查询。它进一步采用主动视角验证机制,判断当前观测是否足够,并在必要时重新定位至更信息丰富的视角,再依据类别与属性约束验证候选目标。尽管专为层次化OVON设计,SAP-Nav无需任务特定训练或预计算场景地图,即可支持层次化和标准类别级OVON。在LangMap和HM3D-OVON数据集上的实验表明,SAP-Nav取得整体最佳性能,尤其在区域级导航上相较训练方法提升12.2%的成功率。真实机器人实验进一步验证了其实际可行性。代码将在录用后公开。

原文摘要 · Abstract (English)

Hierarchical open-vocabulary object navigation (OVON) requires agents to follow free-form instructions that may specify targets through scene-, room-, region-, and instance-level cues in unseen environments. Although recent work LangMap has formalized this setting, reliably solving it under partial observations remains challenging: spatial grounding requires persistent environment-level evidence, whereas target verification requires clear and discriminative candidate views. We present SAP-Nav, a fully online, zero-shot framework that addresses both requirements through active perception. SAP-Nav incrementally constructs a Queryable Spatial-Semantic Representation from actively acquired room views, enabling spatial semantic queries from any explored location. It further employs Active Viewpoint Verification to assess whether the current observation provides sufficient evidence and, when necessary, reposition the agent to a more informative viewpoint before verifying candidates against category and attribute constraints. Although designed for hierarchical OVON, SAP-Nav supports both hierarchical and standard category-level OVON without task-specific training or precomputed scene maps. Experiments on LangMap and HM3D-OVON show that SAP-Nav achieves the overall best performance, including a 12.2% improvement in SR over training-based methods on region-level navigation. Real-world robot experiments further demonstrate its practical feasibility. Code will be made publicly available upon acceptance.

对象导航主动感知空间语义零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。