让AI导航时像人一样预想未来场景,提升路径规划准确率。
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
- 用扩散模型生成基于动作和指令的未来视角图像
- 在线生成价值图使导航成功率在MP3D上提升至82.3%
- 适合研究视觉语言导航与具身智能的开发者
视觉-语言导航(VLN)要求智能体在连续真实环境中根据语言指令行动。以往基于图像想象的方法虽在离散全景图上有优势,但缺乏在线、动作条件化的预测能力,且不生成明确的规划值;许多方法用长时程目标替代规划器,导致结果脆弱且缓慢。为此,我们提出VISTAv2,一种生成式世界模型,可基于过往观测、候选动作序列和语言指令,滚动预测视点视角,并将其投影为在线价值图用于规划。不同于先前方法,VISTAv2不替换原有规划器,而是将在线价值图以分数级融合进基础目标,提供可达性与风险感知引导。具体而言,我们采用动作感知的条件扩散变换器视频预测器生成短时程未来画面,通过视觉-语言评分器对齐自然语言指令,并通过可微的想象到价值头融合多个推演结果,输出想象中的视点价值图。为提升效率,推演在VAE隐空间中进行,使用轻量采样器和稀疏解码,可在单张消费级显卡上完成推理。在MP3D和RoboTHOR上的评估表明,VISTAv2优于强基线,消融实验显示动作条件化想象、指令引导的价值融合及在线价值图规划均至关重要,证明其为鲁棒VLN提供了实用且可解释的路径。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) requires agents to follow language instructions while acting in continuous real-world spaces. Prior image imagination based VLN work shows benefits for discrete panoramas but lacks online, action-conditioned predictions and does not produce explicit planning values; moreover, many methods replace the planner with long-horizon objectives that are brittle and slow. To bridge this gap, we propose VISTAv2, a generative world model that rolls out egocentric future views conditioned on past observations, candidate action sequences, and instructions, and projects them into an online value map for planning. Unlike prior approaches, VISTAv2 does not replace the planner. The online value map is fused at score level with the base objective, providing reachability and risk-aware guidance. Concretely, we employ an action-aware Conditional Diffusion Transformer video predictor to synthesize short-horizon futures, align them with the natural language instruction via a vision-language scorer, and fuse multiple rollouts in a differentiable imagination-to-value head to output an imagined egocentric value map. For efficiency, rollouts occur in VAE latent space with a distilled sampler and sparse decoding, enabling inference on a single consumer GPU. Evaluated on MP3D and RoboTHOR, VISTAv2 improves over strong baselines, and ablations show that action-conditioned imagination, instruction-guided value fusion, and the online value-map planner are all critical, suggesting that VISTAv2 offers a practical and interpretable route to robust VLN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。