用想象引导记忆检索,让智能体更聪明地导航
Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation
- 用世界模型想象未来路径,生成查询来精准找记忆
- 在IR2R上提升5.4%成功率,训练快8.3倍,内存减少74%
- 适合研究长期记忆导航或高效智能体的学者
视觉-语言导航(VLN)要求智能体根据自然语言指令在环境中移动,记忆持久型版本需通过累积经验持续改进。现有方法存在关键缺陷:缺乏有效记忆访问机制,多依赖完整记忆整合或固定视野检索,且仅存储环境观测,忽视了蕴含决策策略的导航行为模式。本文提出Memoir,以想象力为检索机制,基于显式记忆构建:世界模型生成未来状态作为查询,实现对环境观测与行为历史的精准检索。其核心包括:1)语言条件的世界模型,既编码经验又生成检索查询;2)视点级混合记忆,将观测与行为模式锚定于视点,支持混合检索;3)增强型导航模型,通过专用编码器融合检索知识。在包含10种测试场景的多个记忆持久型VLN基准上评估显示,Memoir全面领先,相较最优基线在IR2R上提升5.4% SPL,训练速度提升8.3倍,推理内存降低74%。结果验证了对环境与行为记忆的预测性检索能显著提升导航性能,分析表明该范式仍有73.3%至93.4%的上升空间。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) requires agents to follow natural language instructions through environments, with memory-persistent variants demanding progressive improvement through accumulated experience. Existing approaches for memory-persistent VLN face critical limitations: they lack effective memory access mechanisms, instead relying on entire memory incorporation or fixed-horizon lookup, and predominantly store only environmental observations while neglecting navigation behavioral patterns that encode valuable decision-making strategies. We present Memoir, which employs imagination as a retrieval mechanism grounded by explicit memory: a world model imagines future navigation states as queries to selectively retrieve relevant environmental observations and behavioral histories. The approach comprises: 1) a language-conditioned world model that imagines future states serving dual purposes: encoding experiences for storage and generating retrieval queries; 2) Hybrid Viewpoint-Level Memory that anchors both observations and behavioral patterns to viewpoints, enabling hybrid retrieval; and 3) an experience-augmented navigation model that integrates retrieved knowledge through specialized encoders. Extensive evaluation across diverse memory-persistent VLN benchmarks with 10 distinct testing scenarios demonstrates Memoir's effectiveness: significant improvements across all scenarios, with 5.4% SPL gains on IR2R over the best memory-persistent baseline, accompanied by 8.3x training speedup and 74% inference memory reduction. The results validate that predictive retrieval of both environmental and behavioral memories enables more effective navigation, with analysis indicating substantial headroom (73.3% vs 93.4% upper bound) for this imagination-guided paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。