arXiv:2603.04639cs.ROcs.AI2026-03中稿 · ICML被引 30

构建首个标准化机器人长时记忆评估基准,推动通用机械臂智能进步

RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies

  • 设计16个任务的分类体系,系统评测时间、空间、物体与流程记忆能力
  • 在14种记忆增强模型上验证:不同记忆机制效果因任务而异
  • 适合研究机器人长期决策、记忆机制或具身智能的学者与工程师

记忆对长周期、依赖历史的机器人操作至关重要,此类任务常涉及重复动作计数或临时遮挡物体的操作。近期视觉-语言-动作(VLA)模型开始引入记忆机制,但其评估仍局限于狭窄且非标准化的场景,限制了系统性理解、对比与进展衡量。为解决该问题,我们提出RoboMME:一个大规模标准化基准,用于评估和推进VLA模型在长周期、依赖历史场景中的表现。该基准包含16个基于精心设计分类体系构建的操作任务,涵盖时间、空间、物体与流程记忆的评测。我们进一步基于π0.5骨干网络开发了14种记忆增强型VLA变体,系统探索不同记忆表示在多种集成策略下的表现。实验结果表明,记忆机制的有效性高度依赖任务特性,每种设计在不同任务中均有独特优势与局限。视频与代码见官网https://robomme.github.io。

原文摘要 · Abstract (English)

Memory is critical for long-horizon and history-dependent robotic manipulation. Such tasks often involve counting repeated actions or manipulating objects that become temporarily occluded. Recent vision-language-action (VLA) models have begun to incorporate memory mechanisms; however, their evaluations remain confined to narrow, non-standardized settings. This limits systematic understanding, comparison, and progress measurement. To address these challenges, we introduce RoboMME: a large-scale standardized benchmark for evaluating and advancing VLA models in long-horizon, history-dependent scenarios. Our benchmark comprises 16 manipulation tasks constructed under a carefully designed taxonomy that evaluates temporal, spatial, object, and procedural memory. We further develop a suite of 14 memory-augmented VLA variants built on the π0.5 backbone to systematically explore different memory representations across multiple integration strategies. Experimental results show that the effectiveness of memory representations is highly task-dependent, with each design offering distinct advantages and limitations across different tasks. Videos and code can be found at our website https://robomme.github.io.

机器人长时记忆视觉语言动作基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。