arXiv:2505.07868cs.RO2025-05被引 8

用生成模型想象目标场景,让智能体更准导航

VISTA: Generative Visual Imagination for Vision-and-Language Navigation

  • 用扩散模型根据语言指令和视觉信息动态生成目标图像
  • 在R2R数据集上成功率提升3.6%,刷新基准
  • 适合研究长程视觉语言导航的开发者

视觉-语言导航(VLN)任务要求智能体在未见过的环境中,依据自然语言指令和视觉线索定位特定物体。现有方法多采用‘观察与推理’模式,即基于当前视觉观测决定下一步动作,但在长程场景中受限于即时感知和跨模态差异。为此,本文提出VISTA,一种创新的‘想象与对齐’导航策略。该框架利用预训练扩散模型的生成先验,结合局部观测与高层语言指令进行动态视觉想象;再通过感知对齐过滤模块,将生成的目标图像与当前观测对齐,引导可解释、结构化的推理过程以选择动作。实验表明,VISTA在Room-to-Room(R2R)和RoboTHOR基准上均取得新最佳表现,例如在R2R上成功率达到+3.6%的提升。大量消融实验验证了前瞻性想象、感知对齐与结构化推理对长程环境鲁棒导航的关键作用。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) tasks agents with locating specific objects in unseen environments using natural language instructions and visual cues. Many existing VLN approaches typically follow an 'observe-and-reason' schema, that is, agents observe the environment and decide on the next action to take based on the visual observations of their surroundings. They often face challenges in long-horizon scenarios due to limitations in immediate observation and vision-language modality gaps. To overcome this, we present VISTA, a novel framework that employs an 'imagine-and-align' navigation strategy. Specifically, we leverage the generative prior of pre-trained diffusion models for dynamic visual imagination conditioned on both local observations and high-level language instructions. A Perceptual Alignment Filter module then grounds these goal imaginations against current observations, guiding an interpretable and structured reasoning process for action selection. Experiments show that VISTA sets new state-of-the-art results on Room-to-Room (R2R) and RoboTHOR benchmarks, e.g.,+3.6% increase in Success Rate on R2R. Extensive ablation analysis underscores the value of integrating forward-looking imagination, perceptual alignment, and structured reasoning for robust navigation in long-horizon environments.

视觉导航生成模型语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。