arXiv:2601.10744cs.AIcs.CV2026-01中稿 · CVPR被引 10

提出可长期记忆的智能体探索框架,提升复杂环境下的持续学习能力

Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration

  • 用多模态大模型+强化学习实现主动记忆查询与探索
  • 在长时程任务中显著优于现有方法,提升探索效率与决策质量
  • 适合研究具身智能、长期学习与多任务认知的学者

理想的具身智能体应具备终身学习能力,以应对长周期、复杂任务,实现在通用环境中的持续运行。这不仅要求完成指定任务,还需利用长期情景记忆优化决策。然而,现有主流的一次性具身任务主要关注任务完成结果,忽视了探索过程与记忆利用的重要性。为此,我们提出长期记忆具身探索(LMEE)框架,统一智能体的探索认知与决策行为,推动终身学习。同时构建了相应数据集与基准测试平台LMEE-Bench,包含多目标导航与基于记忆的问答任务,全面评估具身探索的过程与结果。为增强记忆召回与主动探索能力,提出MemoryExplorer方法,通过强化学习微调多模态大语言模型,鼓励主动记忆查询。采用包含动作预测、前沿选择与问答任务的多任务奖励函数,使模型实现主动探索。大量实验表明,该方法在长时程具身任务中显著优于现有先进模型。数据集与代码将公开于https://wangsen99.github.io/papers/lmee/

原文摘要 · Abstract (English)

An ideal embodied agent should possess lifelong learning capabilities to handle long-horizon and complex tasks, enabling continuous operation in general environments. This not only requires the agent to accurately accomplish given tasks but also to leverage long-term episodic memory to optimize decision-making. However, existing mainstream one-shot embodied tasks primarily focus on task completion results, neglecting the crucial process of exploration and memory utilization. To address this, we propose Long-term Memory Embodied Exploration (LMEE), which aims to unify the agent's exploratory cognition and decision-making behaviors to promote lifelong learning. We further construct a corresponding dataset and benchmark, LMEE-Bench, incorporating multi-goal navigation and memory-based question answering to comprehensively evaluate both the process and outcome of embodied exploration. To enhance the agent's memory recall and proactive exploration capabilities, we propose MemoryExplorer, a novel method that fine-tunes a multimodal large language model through reinforcement learning to encourage active memory querying. By incorporating a multi-task reward function that includes action prediction, frontier selection, and question answering, our model achieves proactive exploration. Extensive experiments against state-of-the-art embodied exploration models demonstrate that our approach achieves significant advantages in long-horizon embodied tasks. Our dataset and code will be released at https://wangsen99.github.io/papers/lmee/

具身智能长期记忆强化学习多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。