arXiv:2411.17735cs.CVcs.RO2024-11CVPR被引 95

用多视角快照构建3D场景记忆,让智能体长期探索更聪明。

3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning

  • 用多视角图像作为记忆快照,捕捉已探索区域的丰富视觉信息
  • 引入前沿快照机制,结合已知与潜在新信息做决策
  • 支持增量构建和记忆检索,适合长期自主探索任务

构建紧凑且信息丰富的3D场景表征对有效具身探索与推理至关重要,尤其在复杂环境中的长期任务中。现有方法如基于物体的3D场景图将场景简化为孤立物体及受限文本关系,难以应对需要精细空间理解的查询。同时,这些表征缺乏主动探索与记忆管理的自然机制,限制了其在终身自主系统中的应用。本文提出3D-Mem,一种面向具身智能体的新型3D场景记忆框架。该框架采用信息丰富的多视角图像(称为记忆快照)表示场景,捕获已探索区域的丰富视觉内容;并通过引入前沿快照(未探索区域的瞥见),实现基于已知与潜在新信息的前瞻式探索。为支持主动探索中的终身记忆,我们设计了3D-Mem的增量构建流程及记忆检索技术。在三个基准上的实验表明,3D-Mem显著提升了智能体在3D环境中的探索与推理能力,展现出在具身AI应用中的巨大潜力。

原文摘要 · Abstract (English)

Constructing compact and informative 3D scene representations is essential for effective embodied exploration and reasoning, especially in complex environments over extended periods. Existing representations, such as object-centric 3D scene graphs, oversimplify spatial relationships by modeling scenes as isolated objects with restrictive textual relationships, making it difficult to address queries requiring nuanced spatial understanding. Moreover, these representations lack natural mechanisms for active exploration and memory management, hindering their application to lifelong autonomy. In this work, we propose 3D-Mem, a novel 3D scene memory framework for embodied agents. 3D-Mem employs informative multi-view images, termed Memory Snapshots, to represent the scene and capture rich visual information of explored regions. It further integrates frontier-based exploration by introducing Frontier Snapshots-glimpses of unexplored areas-enabling agents to make informed decisions by considering both known and potential new information. To support lifelong memory in active exploration settings, we present an incremental construction pipeline for 3D-Mem, as well as a memory retrieval technique for memory management. Experimental results on three benchmarks demonstrate that 3D-Mem significantly enhances agents' exploration and reasoning capabilities in 3D environments, highlighting its potential for advancing applications in embodied AI.

3D场景记忆具身智能主动探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。