通过双关系推理提升未知物体导航的泛化能力
Nav-$R^2$ Dual-Relation Reasoning for Generalizable Open-Vocabulary Object-Goal Navigation
- 构建链式思维框架,显式建模目标与环境、环境与动作的关系
- 在未见物体定位任务中达到最新最好性能,实时推理速度2Hz
- 无需额外参数,通过感知记忆压缩保留关键语义特征
开放词汇场景下的物体目标导航要求智能体在未见过的环境中定位新物体,但现有方法存在决策过程不透明、对未见物体定位成功率低的问题。为此,我们提出Nav-$R^2$框架,通过结构化链式思维(CoT)推理和感知记忆(SA-Mem),显式建模目标-环境关系与环境-动作规划两类关键关系。我们构建了Nav$R^2$-CoT数据集,训练模型感知环境、聚焦目标相关物体并制定未来行动策略。SA-Mem通过压缩视频帧并融合历史观测,在时间与语义双重维度上保留最相关的特征,且不引入额外参数。相比先前方法,Nav-$R^2$在定位未见物体上实现最优表现,采用简洁高效流水线,避免对已见类别过拟合,同时保持2Hz实时推理速度。
原文摘要 · Abstract (English)
Object-goal navigation in open-vocabulary settings requires agents to locate novel objects in unseen environments, yet existing approaches suffer from opaque decision-making processes and low success rate on locating unseen objects. To address these challenges, we propose Nav-$R^2$, a framework that explicitly models two critical types of relationships, target-environment modeling and environment-action planning, through structured Chain-of-Thought (CoT) reasoning coupled with a Similarity-Aware Memory. We construct a Nav$R^2$-CoT dataset that teaches the model to perceive the environment, focus on target-related objects in the surrounding context and finally make future action plans. Our SA-Mem preserves the most target-relevant and current observation-relevant features from both temporal and semantic perspectives by compressing video frames and fusing historical observations, while introducing no additional parameters. Compared to previous methods, Nav-R^2 achieves state-of-the-art performance in localizing unseen objects through a streamlined and efficient pipeline, avoiding overfitting to seen object categories while maintaining real-time inference at 2Hz. Resources will be made publicly available at \href{https://github.com/AMAP-EAI/Nav-R2}{github link}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。