arXiv:2507.21450cs.CVcs.RO2025-07

通过递归想象与自适应语义对齐,提升视觉语言导航的指令理解能力。

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation

  • 递归总结沿途视觉信息,聚焦场景规律而非细节干扰
  • 在VLN-CE和ObjectNav上超越当前最优方法,显著提升导航准确率
  • 适合研究视觉语言导航、具身智能与多模态推理的学者

视觉语言导航(VLN)要求智能体根据语言指令在未知场景中导航至目标物体或区域。该任务需对历史视觉观测进行语义对齐,这对长序列决策至关重要。然而现有方法存在场景表征过细、视觉-语言对齐模糊的问题,削弱了对导航友好型高层场景先验的理解,易导致违背指令的行为。为此,我们提出一种递归视觉想象(RVI)策略,将历史轨迹结构化为紧凑神经网格,促使智能体关注视觉变化规律与语义场景布局,而非误导性的几何细节。同时,设计自适应语言接地(ALG)机制,针对性地将情境记忆与不同语言成分对齐,实现细粒度语义匹配,从而精准预测导航动作与进展。该策略在挑战性任务VLN-CE和ObjectNav上优于当前最优方法,验证了RVI与ALG的有效性。

原文摘要 · Abstract (English)

Vision Language Navigation (VLN) typically requires agents to navigate to specified objects or remote regions in unknown scenes by obeying linguistic commands. Such tasks require organizing historical visual observations for linguistic grounding, which is critical for long-sequence navigational decisions. However, current agents suffer from overly detailed scene representation and ambiguous vision-language alignment, which weaken their comprehension of navigation-friendly high-level scene priors and easily lead to behaviors that violate linguistic commands. To tackle these issues, we propose a navigation policy by recursively summarizing along-the-way visual perceptions, which are adaptively aligned with commands to enhance linguistic grounding. In particular, by structurally modeling historical trajectories as compact neural grids, several Recursive Visual Imagination (RVI) techniques are proposed to motivate agents to focus on the regularity of visual transitions and semantic scene layouts, instead of dealing with misleading geometric details. Then, an Adaptive Linguistic Grounding (ALG) technique is proposed to align the learned situational memories with different linguistic components purposefully. Such fine-grained semantic matching facilitates the accurate anticipation of navigation actions and progress. Our navigation policy outperforms the state-of-the-art methods on the challenging VLN-CE and ObjectNav tasks, showing the superiority of our RVI and ALG techniques for VLN.

视觉语言导航递归想象语义对齐具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。