arXiv:2603.18443cs.CV2026-03

利用物体间空间关系提升零样本目标导航的鲁棒性

SR-Nav: Spatial Relationships Matter for Zero-shot Object Goal Navigation

  • 构建动态空间关系图,融合观测与先验关系
  • 在HM3D上成功率达87.3%,路径效率提升32%
  • 适合需要强泛化能力的视觉导航研究者

零样本物体目标导航旨在仅通过自身视角观察,在未见过的环境中找到目标物体。现有方法依赖基础模型的理解与推理能力,但在视角不佳或语义线索弱时,其感知与规划易出错。我们发现物体与区域间的内在关系蕴含结构化场景先验,能帮助智能体在部分观测下推断合理的目标位置。为此,提出空间关系感知导航(SR-Nav)框架,通过建模实时观测与经验积累的空间关系,增强感知与规划。首先构建动态空间关系图(DSRG),以目标为中心编码空间关系并随观测动态更新;其次引入关系匹配模块,通过关系比对而非直接检测来验证和纠正误判,提升视觉感知鲁棒性;最后设计动态关系规划模块,基于当前位姿从DSRG中计算最优路径,减少搜索空间与探索冗余。在HM3D数据集上的实验表明,本方法在成功率与导航效率上均达到当前最佳水平。

原文摘要 · Abstract (English)

Zero-shot object-goal navigation aims to find target objects in unseen environments using only egocentric observation. Recent methods leverage foundation models' comprehension and reasoning capabilities to enhance navigation performance. However, when faced with poor viewpoints or weak semantic cues, foundation models often fail to support reliable reasoning in both perception and planning, resulting in inefficient or failed navigation. We observe that inherent relationships among objects and regions encode structured scene priors, which help agents infer plausible target locations even under partial observations. Motivated by this insight, we propose Spatial Relation-aware Navigation (SR-Nav), a framework that models both observed and experience-based spatial relationships to enhance both perception and planning. Specifically, SR-Nav first constructs a Dynamic Spatial Relationship Graph (DSRG) that encodes the target-centered spatial relationships through the foundation models and updates dynamically with real-time observations. We then introduce a Relation-aware Matching Module. It utilizes relationship matching instead of naive detection, leveraging diverse relationships in the DSRG to verify and correct errors, enhancing visual perception robustness. Finally, we design a Dynamic Relationship Planning Module to reduce the planning search space by dynamically computing the optimal paths based on the DSRG from the current position, thereby guiding planning and reducing exploration redundancy. Experiments on HM3D show that our method achieves state-of-the-art performance in both success rate and navigation efficiency. The code will be publicly available at https://github.com/Mzyw-1314/SR-Nav

导航空间关系零样本视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。