让机器人主动选择记忆什么,提升长时操作能力。
GaussMemory: Task-Driven 3D Gaussian Scene Memory for Long-Horizon Robotic Manipulation

- 基于3D高斯溅射构建可动态更新的记忆体,任务驱动记忆策略。
- 在LIBERO和VLABench上分别超越基线5.2%~6.0%,显著提升成功率。
- 适合研究长时机器人操作与具身智能的开发者参考。
长时程机器人操作依赖持久的空间记忆。然而现有3D记忆系统仅作为被动记录器:使用固定的手工规则存储观察结果,对所有场景元素——无论是否为关键抓取目标或无关背景墙——一视同仁。本文提出从被动存储到主动、任务驱动空间记忆的范式转变。我们认为,机器人的记忆不应仅记录所见,而应主动学习如何记忆——自主发现应追踪哪些物体、以多强力度更新以及何时丢弃,全部通过端到端学习实现,无需手工设计规则。关键在于,该主动范式通过将记忆更新与读取统一为同一认知过程的两个方面实现,支持双向流动:任务需求塑造更新策略,反之亦然。为实现这一愿景,我们引入GaussMemory,利用3D高斯溅射作为持久几何基础。在LIBERO上,GaussMemory在目标和长序列10任务上优于MemoryVLA;在VLABench上,相比$π_0$-FAST分别提升+5.2%(跟踪1)和+6.0%(跟踪6)。
原文摘要 · Abstract (English)
Long-horizon robotic manipulation fundamentally relies on persistent spatial memory. However, existing 3D memory systems function merely as passive recorders: they store observations using fixed, hand-crafted rules, treating every scene element--whether a critical grasp target or an irrelevant background wall--with equal importance. In this paper, we propose a paradigm shift from passive storage to active, task-driven spatial memory. We argue that a robot's memory should not simply record what it sees, but actively learn how to remember--discovering which objects to track precisely, how aggressively to update them, and what to discard, all learned end-to-end without hand-designed rules. Crucially, this active paradigm is realized by unifying memory update and readout as two sides of the same cognitive process, enabling bidirectional flow where task needs shape update strategies and vice versa. To instantiate this vision, we introduce GaussMemory, which leverages 3D Gaussian Splatting as a persistent geometric substrate. On LIBERO, GaussMemory outperforms MemoryVLA on Goal and Long-10; on VLABench, it surpasses $π_0$-FAST by +5.2% (Track 1) and +6.0% (Track 6).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。