arXiv:2509.25655cs.AI2025-09中稿 · publication by Int…

用外部知识库+地标引导,提升导航模型理解指令能力。

Landmark-Guided Knowledge for Vision-and-Language Navigation

  • 构建63万条描述的知识库,通过视觉匹配提取相关知识。
  • 基于地标信息引导注意力,降低外部知识引入的偏差。
  • 融合语言、视觉与历史信息,显著提升导航准确率和效率。

视觉-语言导航是具身智能的核心任务,要求智能体根据自然语言指令在陌生环境中自主导航。现有方法在复杂场景中常因缺乏常识推理能力而无法准确匹配指令与环境信息。本文提出一种名为地标引导知识(LGK)的导航方法,引入外部知识库辅助决策,解决传统方法因常识不足导致的误判问题。首先,构建包含63万条语言描述的知识库,并通过知识匹配将环境子视图与知识库对齐,提取相关描述性知识;其次,设计地标引导的知识机制(KGL),利用指令中的地标信息引导智能体关注最相关知识,减少引入外部知识带来的数据偏差;最后,提出知识引导的动态增强(KGDA)策略,有效融合语言、知识、视觉及历史信息。实验结果表明,LGK在R2R和REVERIE数据集上优于现有最优方法,尤其在导航误差、成功率和路径效率方面表现突出。

原文摘要 · Abstract (English)

Vision-and-language navigation is one of the core tasks in embodied intelligence, requiring an agent to autonomously navigate in an unfamiliar environment based on natural language instructions. However, existing methods often fail to match instructions with environmental information in complex scenarios, one reason being the lack of common-sense reasoning ability. This paper proposes a vision-and-language navigation method called Landmark-Guided Knowledge (LGK), which introduces an external knowledge base to assist navigation, addressing the misjudgment issues caused by insufficient common sense in traditional methods. Specifically, we first construct a knowledge base containing 630,000 language descriptions and use knowledge Matching to align environmental subviews with the knowledge base, extracting relevant descriptive knowledge. Next, we design a Knowledge-Guided by Landmark (KGL) mechanism, which guides the agent to focus on the most relevant parts of the knowledge by leveraging landmark information in the instructions, thereby reducing the data bias that may arise from incorporating external knowledge. Finally, we propose Knowledge-Guided Dynamic Augmentation (KGDA), which effectively integrates language, knowledge, vision, and historical information. Experimental results demonstrate that the LGK method outperforms existing state-of-the-art methods on the R2R and REVERIE vision-and-language navigation datasets, particularly in terms of navigation error, success rate, and path efficiency.

视觉导航知识增强语言理解具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。