通过推理目标与视角的空间关系,提升图像导航的精准度。
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
- 构建目标与当前视角的关联关系作为导航引导
- 在三个基准数据集上表现更优,尤其在用户匹配场景
- 适合需要精准导航的真实场景应用
近期图像目标导航(ImageNav)方法通过分别提取目标和自身视角图像的语义特征,再输入策略网络生成动作。然而仍面临两大挑战:(1) 语义特征难以提供准确方向信息,导致冗余动作;(2) 当训练与实际应用存在视角不一致时,性能显著下降。为此,我们提出RSRNav,一种简单而有效的方法,通过推理目标与当前观测之间的空间关系来指导导航。具体而言,我们通过构建目标与当前观测间的相关性来建模空间关系,并将其传递至策略网络进行动作预测。这些相关性通过细粒度交叉相关和方向感知相关逐步优化,实现更精确导航。在三个基准数据集上的广泛评估表明,RSRNav在导航性能上表现优异,尤其在“用户匹配目标”设置下,凸显其在真实场景中的应用潜力。
原文摘要 · Abstract (English)
Recent image-goal navigation (ImageNav) methods learn a perception-action policy by separately capturing semantic features of the goal and egocentric images, then passing them to a policy network. However, challenges remain: (1) Semantic features often fail to provide accurate directional information, leading to superfluous actions, and (2) performance drops significantly when viewpoint inconsistencies arise between training and application. To address these challenges, we propose RSRNav, a simple yet effective method that reasons spatial relationships between the goal and current observations as navigation guidance. Specifically, we model the spatial relationship by constructing correlations between the goal and current observations, which are then passed to the policy network for action prediction. These correlations are progressively refined using fine-grained cross-correlation and direction-aware correlation for more precise navigation. Extensive evaluation of RSRNav on three benchmark datasets demonstrates superior navigation performance, particularly in the "user-matched goal" setting, highlighting its potential for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。