arXiv:2603.26859cs.CVcs.AI2026-03中稿 · publication by the…

用图文知识库提升视觉语言导航的语义对齐能力

Beyond Textual Knowledge-Leveraging Multimodal Knowledge Bases for Enhancing Vision-and-Language Navigation

  • 融合环境文本与生成图像知识库,增强语义理解
  • 在R2R和REVERIE上准确率提升5%和2.07%
  • 适合做多模态导航、智能机器人路径规划的研究者

视觉-语言导航(VLN)要求智能体根据自然语言指令在复杂未见环境中导航。现有方法常难以有效捕捉关键语义线索并精确对齐视觉观察。为此,本文提出超越文本知识(BTK)框架,协同整合环境特异性文本知识与生成式图像知识库。BTK使用Qwen3-4B提取目标相关短语,利用Flux-Schnell构建两个大规模图像知识库:R2R-GP与REVERIE-GP。同时,通过BLIP-2从全景视图构建大规模文本知识库,提供环境特异性语义线索。这些多模态知识库通过目标感知增强器与知识增强器有效融合,显著提升语义定位与跨模态对齐能力。在包含7,189条轨迹的R2R数据集和21,702条指令的REVERIE数据集上进行的大量实验表明,BTK显著优于现有基线。在R2R和REVERIE的测试未见划分上,成功率(SR)分别提升5%和2.07%,路径相似度(SPL)分别提升4%和3.69%。源代码已开源。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) requires an agent to navigate through complex unseen environments based on natural language instructions. However, existing methods often struggle to effectively capture key semantic cues and accurately align them with visual observations. To address this limitation, we propose Beyond Textual Knowledge (BTK), a VLN framework that synergistically integrates environment-specific textual knowledge with generative image knowledge bases. BTK employs Qwen3-4B to extract goal-related phrases and utilizes Flux-Schnell to construct two large-scale image knowledge bases: R2R-GP and REVERIE-GP. Additionally, we leverage BLIP-2 to construct a large-scale textual knowledge base derived from panoramic views, providing environment-specific semantic cues. These multimodal knowledge bases are effectively integrated via the Goal-Aware Augmentor and Knowledge Augmentor, significantly enhancing semantic grounding and cross-modal alignment. Extensive experiments on the R2R dataset with 7,189 trajectories and the REVERIE dataset with 21,702 instructions demonstrate that BTK significantly outperforms existing baselines. On the test unseen splits of R2R and REVERIE, SR increased by 5% and 2.07% respectively, and SPL increased by 4% and 3.69% respectively. The source code is available at https://github.com/yds3/IPM-BTK/.

视觉导航多模态知识库语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。