arXiv:2412.01857cs.CVcs.LG2024-12AAAI被引 12

让AI像人一样想象未来场景,提升导航能力。

Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation

  • 构建现实与想象结合的记忆系统,支持动态记忆更新。
  • 能生成未来场景的高保真图像,SPL指标达最新最佳水平。
  • 适合研究具身智能、视觉语言导航及想象力建模的读者。

人类在陌生环境中导航时依赖情景模拟与情景记忆,有助于深入理解环境与物体间的复杂关系。受此启发,我们提出一种新型架构,为智能体配备现实-想象混合记忆系统,使其通过想象机制和导航行为共同维护并扩展记忆。此外,设计了针对性预训练任务以增强智能体的想象能力。该智能体可生成未来场景的高保真RGB图像,在未见过的环境中显著提升导航表现,成功实现视觉-语言导航(VLN)任务中路径长度加权成功率(SPL)的当前最优结果。

原文摘要 · Abstract (English)

Humans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memory system inspired by human mechanisms can enhance the navigation performance of embodied agents in unseen environments. However, existing Vision-and-Language Navigation (VLN) agents lack a memory mechanism of this kind. To address this, we propose a novel architecture that equips agents with a reality-imagination hybrid memory system. This system enables agents to maintain and expand their memory through both imaginative mechanisms and navigation actions. Additionally, we design tailored pre-training tasks to develop the agent's imaginative capabilities. Our agent can imagine high-fidelity RGB images for future scenes, achieving state-of-the-art result in Success rate weighted by Path Length (SPL).

视觉导航情景记忆想象力建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。