arXiv:2509.25139cs.AIcs.CV2025-09EMNLP被引 4

用多视角文本描述提升大模型导航的抽象推理能力

Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs

  • 引入多角度文本描述,支持跨图像类比推理
  • 在R2R数据集上导航成功率显著提升
  • 适合研究视觉语言导航与大模型推理融合的学者

将大语言模型(LLM)融入具身智能系统日益普遍。现有零样本基于LLM的视觉-语言导航(VLN)代理要么将图像编码为文本场景描述,可能简化视觉细节;要么处理原始图像输入,难以捕捉高层推理所需的抽象语义。本文通过引入多视角文本描述,促进图像间的类比推理,增强代理对全局场景和空间关系的理解,从而做出更准确的动作决策。我们在R2R数据集上验证了该方法,实验表明导航性能显著提升。

原文摘要 · Abstract (English)

Integrating large language models (LLMs) into embodied AI models is becoming increasingly prevalent. However, existing zero-shot LLM-based Vision-and-Language Navigation (VLN) agents either encode images as textual scene descriptions, potentially oversimplifying visual details, or process raw image inputs, which can fail to capture abstract semantics required for high-level reasoning. In this paper, we improve the navigation agent's contextual understanding by incorporating textual descriptions from multiple perspectives that facilitate analogical reasoning across images. By leveraging text-based analogical reasoning, the agent enhances its global scene understanding and spatial reasoning, leading to more accurate action decisions. We evaluate our approach on the R2R dataset, where our experiments demonstrate significant improvements in navigation performance.

视觉语言导航大模型类比推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。