arXiv:2503.11006cs.CVcs.AI2025-03被引 2

通过细粒度指令引导图推理,提升视觉语言导航的精准度。

Fine-Grained Instruction-Guided Graph Reasoning for Vision-and-Language Navigation

  • 分离视觉与方向信息,用几何嵌入强化空间图推理。
  • 从指令中提取位置和物体关键语义,实现更准跨模态对齐。
  • 适合研究视觉语言导航、多模态推理的学者参考。

视觉-语言导航(VLN)要求智能体根据自然语言指令穿越复杂环境,需在视觉观测与语言指引间建立精确对齐。现有方法通常耦合编码视觉与方向线索,并未显式提取导航关键语义,导致空间推理不准确、跨模态对齐不佳。为此,我们提出细粒度指令引导图推理框架(OIKG),增强导航过程中的空间表征与指令理解。具体地,引入观察图交互机制,解耦角度与视觉线索,通过几何嵌入强化有向边表示,提升导航图内的空间推理可靠性;同时设计细粒度指令引导模块,显式提取语言指令中的位置特异性与物体中心信息,促进语言语义与可导航轨迹间的精准对齐。联合结构化图推理与指令关键语义线索,显著提升智能体遵循复杂导航指令的能力。在R2R和RxR基准上的大量实验表明,该方法在多个评估指标上持续达到最先进性能,验证了细粒度指令引导图推理在视觉-语言导航中的有效性。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) requires an embodied agent to traverse complex environments by following natural language instructions, demanding accurate alignment between visual observations and linguistic guidance. Despite recent progress, existing methods typically encode visual and directional cues in a coupled manner, and process instructions without explicitly extracting navigation-critical semantics, which often leads to imprecise spatial reasoning and suboptimal cross-modal alignment. To address these challenges, we propose a fine-grained instruction-guided graph reasoning framework (OIKG) that enhances both spatial representation and instruction understanding during navigation. Specifically, an observation-graph interaction mechanism is introduced to disentangle angular and visual cues while strengthening directed edge representations through geometric embedding, enabling more reliable spatial reasoning within the navigation graph. In addition, a fine-grained instruction guidance module is designed to explicitly extract and leverage location-specific and object-centric information from language instructions, facilitating more precise cross-modal alignment between linguistic semantics and navigable trajectories. By jointly integrating structured graph reasoning with instruction-critical semantic cues, the proposed approach significantly improves the agent's ability to follow complex navigation instructions. Extensive experiments on the R2R and RxR benchmarks demonstrate that our method consistently achieves state-of-the-art performance across multiple evaluation metrics, validating the effectiveness of fine-grained instruction-guided graph reasoning for vision-and-language navigation.

视觉语言导航图神经网络指令理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。