arXiv:2603.03596cs.ROcs.LG2026-03被引 58

让机器人同时记住长期任务目标和短期操作细节,提升复杂任务执行能力。

MEM: Multi-Scale Embodied Memory for Vision Language Action Models

  • 用视频+文本混合记忆,分层次记录短期动作与长期任务进展
  • 支持长达十五分钟的连续任务,如清洁厨房或做三明治
  • 能根据记忆动态调整操作策略,适合长时序机器人控制场景

传统端到端机器人学习中的记忆仅将历史观察序列输入策略模型。但在复杂的多阶段现实任务中,机器人需在不同抽象层级上表征过去事件:从捕捉抽象语义概念的长期记忆(如烹饪晚餐时记住已完成的菜谱步骤),到补偿遮挡的短期记忆(如手臂遮挡后仍记得要抓取的物体)。本文核心洞察是,有效的长时序机器人控制记忆架构应融合多种模态。我们提出多尺度具身记忆(MEM),结合视频编码器压缩的短时记忆与文本驱动的长时记忆,实现混合模态的长时记忆。该方法使机器人策略可完成长达十五分钟的任务,如清理厨房或制作烤奶酪三明治。此外,实验发现记忆使策略能在上下文中智能调整操作策略。

原文摘要 · Abstract (English)

Conventionally, memory in end-to-end robotic learning involves inputting a sequence of past observations into the learned policy. However, in complex multi-stage real-world tasks, the robot's memory must represent past events at multiple levels of granularity: from long-term memory that captures abstracted semantic concepts (e.g., a robot cooking dinner should remember which stages of the recipe are already done) to short-term memory that captures recent events and compensates for occlusions (e.g., a robot remembering the object it wants to pick up once its arm occludes it). In this work, our main insight is that an effective memory architecture for long-horizon robotic control should combine multiple modalities to capture these different levels of abstraction. We introduce Multi-Scale Embodied Memory (MEM), an approach for mixed-modal long-horizon memory in robot policies. MEM combines video-based short-horizon memory, compressed via a video encoder, with text-based long-horizon memory. Together, they enable robot policies to perform tasks that span up to fifteen minutes, like cleaning up a kitchen, or preparing a grilled cheese sandwich. Additionally, we find that memory enables MEM policies to intelligently adapt manipulation strategies in-context.

机器人控制多模态记忆长时序任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。