arXiv:2603.03745cs.AIcs.RO2026-03被引 1

用拓扑结构增强多目标视觉语言导航的推理能力

RAGNav: A Retrieval-Augmented Topological Reasoning Framework for Multi-Goal Visual-Language Navigation

  • 构建双基记忆系统,融合拓扑地图与语义森林
  • 通过邻域得分传播提升目标可达性推理效率
  • 适合复杂场景下多目标路径规划的研究者

视觉语言导航(VLN)正从单目标路径规划向更具挑战性的多目标VLN演进。该任务要求智能体在协同推理空间物理约束与执行顺序的同时,准确识别多个目标实体。然而,通用检索增强生成(RAG)框架在处理多对象关联时,因缺乏显式空间建模,常出现空间幻觉和规划漂移。为此,我们提出RAGNav框架,弥合语义推理与物理结构之间的鸿沟。其核心是双基记忆系统,结合低层拓扑地图以维护物理连通性,以及高层语义森林实现环境层次抽象。基于此表示,框架引入锚点引导的条件检索与拓扑邻域得分传播机制,可快速筛选候选目标、消除语义噪声,并通过拓扑邻域中的物理关联进行语义校准。该机制显著提升了目标间可达性推理能力与序列规划效率。实验表明,RAGNav在复杂多目标导航任务中达到当前最优性能。

原文摘要 · Abstract (English)

Vision-Language Navigation (VLN) is evolving from single-point pathfinding toward the more challenging Multi-Goal VLN. This task requires agents to accurately identify multiple entities while collaboratively reasoning over their spatial-physical constraints and sequential execution order. However, generic Retrieval-Augmented Generation (RAG) paradigms often suffer from spatial hallucinations and planning drift when handling multi-object associations due to the lack of explicit spatial modeling.To address these challenges, we propose RAGNav, a framework that bridges the gap between semantic reasoning and physical structure. The core of RAGNav is a Dual-Basis Memory system, which integrates a low-level topological map for maintaining physical connectivity with a high-level semantic forest for hierarchical environment abstraction. Building on this representation, the framework introduces an anchor-guided conditional retrieval and a topological neighbor score propagation mechanism. This approach facilitates the rapid screening of candidate targets and the elimination of semantic noise, while performing semantic calibration by leveraging the physical associations inherent in the topological neighborhood.This mechanism significantly enhances the capability of inter-target reachability reasoning and the efficiency of sequential planning. Experimental results demonstrate that RAGNav achieves state-of-the-art (SOTA) performance in complex multi-goal navigation tasks.

视觉语言导航多目标推理拓扑建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。