构建机器人操作记忆评估基准,揭示模型设计对记忆能力的影响
RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design
- 提出9个分层级的记忆依赖任务,系统评估策略记忆能力
- 通过实验证明现有策略在长期信息保持上存在明显短板
- 适合研究机器人决策、强化学习与记忆机制的学者参考
近年来,机器人操作策略发展迅速,但多数方法忽视了记忆能力的重要性。这导致它们难以完成需要基于历史观测进行推理或长时间维持任务相关信息的任务,而这类需求在真实操作场景中极为常见。尽管已有部分具备记忆功能的策略被提出,但对记忆依赖性操作的系统评估仍不充分,架构设计选择与记忆性能之间的关系也尚未明确。为此,我们提出了RMBench,一个包含9个操作任务的仿真基准,涵盖不同复杂度的记忆需求,支持对策略记忆能力的系统评估。我们还设计了Mem-0,一种具有显式记忆组件的模块化操作策略,便于开展可控的消融研究。通过大量仿真和真实世界实验,我们揭示了现有策略在记忆方面的局限性,并提供了关于架构设计如何影响记忆性能的实证洞察。
原文摘要 · Abstract (English)
Robotic manipulation policies have made rapid progress in recent years, yet most existing approaches give limited consideration to memory capabilities. Consequently, they struggle to solve tasks that require reasoning over historical observations and maintaining task-relevant information over time, which are common requirements in real-world manipulation scenarios. Although several memory-aware policies have been proposed, systematic evaluation of memory-dependent manipulation remains underexplored, and the relationship between architectural design choices and memory performance is still not well understood. To address this gap, we introduce RMBench, a simulation benchmark comprising 9 manipulation tasks that span multiple levels of memory complexity, enabling systematic evaluation of policy memory capabilities. We further propose Mem-0, a modular manipulation policy with explicit memory components designed to support controlled ablation studies. Through extensive simulation and real-world experiments, we identify memory-related limitations in existing policies and provide empirical insights into how architectural design choices influence memory performance. The website is available at https://rmbench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。