为强化学习机器人记忆能力设计首个综合评估基准
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
- 构建记忆型强化学习任务分类框架,统一评估标准
- 推出32个桌面操作任务,测试机器人在部分可观测下的记忆表现
- 适合研究具身智能、机器人决策与长期记忆的学者使用
记忆对解决具有时空依赖性的复杂任务至关重要。尽管许多强化学习算法引入了记忆机制,但领域内缺乏跨场景评估智能体记忆能力的通用基准。这一缺口在桌面机器人操作中尤为突出——此类任务依赖记忆应对部分可观测性并保障鲁棒性,却无标准化评测体系。为此,我们提出MIKASA(Memory-Intensive Skills Assessment Suite for Agents),包含三项核心贡献:(1) 构建记忆密集型强化学习任务的系统分类框架;(2) 收集MIKASA-Base,实现多场景下记忆增强智能体的系统化评估;(3) 开发MIKASA-Robo(pip install mikasa-robo-suite),包含32个精心设计的记忆密集型任务,专门评估桌面机器人操作中的记忆能力。本工作提供统一框架,推动记忆强化学习研究,助力真实世界系统建设。MIKASA已开源:https://tinyurl.com/membenchrobots。
原文摘要 · Abstract (English)
Memory is crucial for enabling agents to tackle complex tasks with temporal and spatial dependencies. While many reinforcement learning (RL) algorithms incorporate memory, the field lacks a universal benchmark to assess an agent's memory capabilities across diverse scenarios. This gap is particularly evident in tabletop robotic manipulation, where memory is essential for solving tasks with partial observability and ensuring robust performance, yet no standardized benchmarks exist. To address this, we introduce MIKASA (Memory-Intensive Skills Assessment Suite for Agents), a comprehensive benchmark for memory RL, with three key contributions: (1) we propose a comprehensive classification framework for memory-intensive RL tasks, (2) we collect MIKASA-Base -- a unified benchmark that enables systematic evaluation of memory-enhanced agents across diverse scenarios, and (3) we develop MIKASA-Robo (pip install mikasa-robo-suite) -- a novel benchmark of 32 carefully designed memory-intensive tasks that assess memory capabilities in tabletop robotic manipulation. Our work introduces a unified framework to advance memory RL research, enabling more robust systems for real-world use. MIKASA is available at https://tinyurl.com/membenchrobots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。