为3D智能体设计长时空间记忆,让大模型像人一样记住走过的路。
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
- 用动态记忆融合机制,从过往经验中挑选关键信息
- 在复杂3D任务中成功率达89.4%,比最强基线高16.5%
- 适合研究具身智能、长期记忆与多房间导航的学者
人类能通过跨时间和空间的经验记忆完成复杂任务,而当前大语言模型在动态多房间3D环境中难以有效规划与行动。我们认为这一局限部分源于缺乏对3D时空记忆的建模。为此,我们首先提出3DMem-Bench基准,包含超过26,000条轨迹和2,892个具身任务、问答与图像描述任务,用于评估智能体在3D环境中进行长时记忆推理的能力。其次,我们提出3DLLM-Mem,一种面向具身智能体的动态记忆管理与融合模型。该模型使用工作记忆项作为查询,有选择性地关注并融合来自情景记忆中的空间与时间特征,情景记忆存储过去的观察与交互。该方法使智能体能在复杂长时程环境中聚焦任务相关信息的同时保持内存效率。实验表明,3DLLM-Mem在3DMem-Bench最具有挑战性的野外具身任务上,成功率达到89.4%,比最强基线提升16.5%。
原文摘要 · Abstract (English)
Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments. We posit that part of this limitation is due to the lack of proper 3D spatial-temporal memory modeling in LLMs. To address this, we first introduce 3DMem-Bench, a comprehensive benchmark comprising over 26,000 trajectories and 2,892 embodied tasks, question-answering and captioning, designed to evaluate an agent's ability to reason over long-term memory in 3D environments. Second, we propose 3DLLM-Mem, a novel dynamic memory management and fusion model for embodied spatial-temporal reasoning and actions in LLMs. Our model uses working memory tokens, which represents current observations, as queries to selectively attend to and fuse the most useful spatial and temporal features from episodic memory, which stores past observations and interactions. Our approach allows the agent to focus on task-relevant information while maintaining memory efficiency in complex, long-horizon environments. Experimental results demonstrate that 3DLLM-Mem achieves state-of-the-art performance across various tasks, outperforming the strongest baselines by 16.5% in success rate on 3DMem-Bench's most challenging in-the-wild embodied tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。