arXiv:2605.10921cs.RO2026-05被引 11

构建首个真实世界与仿真结合的机器人记忆评测基准,推动长时序任务研究。

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

论文配图:RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
图 1 · 摘自论文原文
  • 用视觉语言模型生成26个复杂任务,平均轨迹超1000步,68.9%任务依赖记忆。
  • 提出PrediMem双系统架构,在真实与仿真环境中显著优于基线模型。
  • 适合关注机器人长期记忆、具身智能和真实场景评估的研究者。

记忆是机器人智能的核心,因其需在部分可观测环境下依赖过往观测与动作完成长时程任务。然而现有机器人记忆评测基准仍缺乏多模态记忆标注,任务覆盖与结构复杂度有限,且仅限于仿真环境。为此,我们提出RoboMemArena,一个包含26个任务的大规模基准,每个任务平均轨迹长度超过1000步,68.9%的子任务为记忆依赖型。其生成流程利用视觉语言模型(VLM)设计并组合子任务,通过原子函数生成完整轨迹,并提供子任务指令与关键帧标注;同时配备真实世界记忆任务以支持物理评估。我们进一步设计PrediMem,一种双系统视觉-语言-动作模型(VLA),其中高层VLM规划器管理近期与关键帧缓冲区记忆库,并采用预测编码头提升对任务动态的敏感性。在RoboMemArena上的大量实验表明,PrediMem超越所有基线,揭示了记忆管理、模型架构及复杂记忆系统的扩展规律。

原文摘要 · Abstract (English)

Memory is a critical component of robotic intelligence, as robots must rely on past observations and actions to accomplish long-horizon tasks in partially observable environments. However, existing robotic memory benchmarks still lack multimodal annotations for memory formation, provide limited task coverage and structural complexity, and remain restricted to simulation without real-world evaluation. We address this gap with RoboMemArena, a large-scale benchmark of 26 tasks, with average trajectory lengths exceeding 1,000 steps per task and 68.9% of subtasks being memory-dependent. The generation pipeline leverages a vision-language model (VLM) to design and compose subtasks, generates full trajectories through atomic functions, and provides memory-related annotations, including subtask instructions and native keyframe annotations, while paired real-world memory tasks support physical evaluation. We further design PrediMem, a dual-system VLA in which a high-level VLM planner manages a memory bank with recent and keyframe buffers and uses a predictive coding head to improve sensitivity to task dynamics. Extensive experiments on RoboMemArena show that PrediMem outperforms all baselines and provides insights into memory management, model architecture, and scaling laws for complex memory systems.

机器人记忆长时序任务视觉语言模型真实世界评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。