arXiv:2505.00743cs.CVcs.RO2025-05中稿 · publication by ICM…被引 10

通过双路物体感知增强,提升视觉语言导航中对指令与环境的细粒度理解。

DOPE: Dual Object Perception-Enhancement Network for Vision-and-Language Navigation

  • 从指令中提取关键语义短语,结合文本与图像中的物体信息进行联合建模。
  • 在R2R和REVERIE数据集上,导航成功率分别提升3.2%和2.8%。
  • 适合需要精准理解语言指令与物体关系的智能机器人导航场景。

视觉-语言导航(VLN)要求智能体根据语言指令,在陌生环境中利用视觉线索完成定位与导航任务。现有方法存在两大局限:其一,直接将完整指令输入多层Transformer网络,未能充分挖掘指令中的细节信息,限制了语言理解能力;其二,忽略跨模态物体间关系的建模,难以利用对象间的潜在关联,影响导航决策的准确性与鲁棒性。为此,我们提出双物体感知增强网络(DOPE)。首先设计文本语义提取模块(TSE),从指令中提取关键短语,并输入文本物体感知增强模块(TOPA),以充分挖掘其中的对象与动作信息;其次引入图像物体感知增强模块(IOPA),对跨模态物体信息进行额外建模,使模型更有效地利用图像与文本间的潜在关联,提升决策精度。在R2R与REVERIE数据集上的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) is a challenging task where an agent must understand language instructions and navigate unfamiliar environments using visual cues. The agent must accurately locate the target based on visual information from the environment and complete tasks through interaction with the surroundings. Despite significant advancements in this field, two major limitations persist: (1) Many existing methods input complete language instructions directly into multi-layer Transformer networks without fully exploiting the detailed information within the instructions, thereby limiting the agent's language understanding capabilities during task execution; (2) Current approaches often overlook the modeling of object relationships across different modalities, failing to effectively utilize latent clues between objects, which affects the accuracy and robustness of navigation decisions. We propose a Dual Object Perception-Enhancement Network (DOPE) to address these issues to improve navigation performance. First, we design a Text Semantic Extraction (TSE) to extract relatively essential phrases from the text and input them into the Text Object Perception-Augmentation (TOPA) to fully leverage details such as objects and actions within the instructions. Second, we introduce an Image Object Perception-Augmentation (IOPA), which performs additional modeling of object information across different modalities, enabling the model to more effectively utilize latent clues between objects in images and text, enhancing decision-making accuracy. Extensive experiments on the R2R and REVERIE datasets validate the efficacy of the proposed approach.

视觉语言导航物体关系多模态理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。