arXiv:2607.14252cs.ROcs.AI2026-07中稿 · EMNLP

让机器人从第一视角视频中构建可长期使用的记忆,支持未来推理与规划。

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

论文配图:MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
图 1 · 摘自论文原文
  • 通过形成-巩固-检索生命周期,将经验组织为多存储记忆结构。
  • 在45小时第一视角视频上,规划准确率提升最高达16.6%。
  • 适合需要个性化记忆的机器人交互与长时任务规划场景。

具身智能体随时间积累经验。本文研究如何将这些经验转化为持久记忆以支持未来的推理与行动。提出具身动作记忆(EAM)概念,即基于具身经验形成并利用记忆的能力,以及由此产生的持久记忆状态。引入MEMORA框架,通过形成-巩固-检索生命周期和多存储世界记忆架构实现EAM。MEMORA将经验组织为个体特定的环境、实体、活动与推断知识存储:在线编辑根据新证据修正记忆,离线巩固则将重复经验抽象为可复用的常规、习惯与偏好。在包含45小时第一视角视频的MEMORA-Bench上评估,该框架在四个开放权重回答模型中取得最强综合规划表现,尤其在分布外规划任务中提升显著。在这些任务中,MEMORA使机器人接地计划得分最高提升16.6%,表明跨经验形成的记忆可支撑对未直接观察目标的新规划。物理机器人演示进一步证明,仅基于人类第一视角视频构建的记忆即可将高层机器人计划锚定于个体特定对象与偏好。

原文摘要 · Abstract (English)

Embodied agents accumulate experience over time. We study how accumulated experience can be formed into persistent memory for future reasoning and action. We formulate Embodied Action Memory (EAM) as the capability to form and use memory over embodied experience, together with the persistent memory state produced by that process. We introduce MEMORA, a framework that instantiates EAM through a formation-consolidation-retrieval lifecycle and a multi-store world-memory architecture. MEMORA organizes experience into participant-specific Environment, Entity, Activity, and Inferred Knowledge stores: online editing revises memory as new evidence arrives, while offline consolidation abstracts repeated experience into reusable routines, habits, and preferences. We evaluate MEMORA with MEMORA-Bench, a 45-hour egocentric-video suite that measures both retrospective memory faithfulness and prospective memory-grounded planning. Across four open-weight answer models, MEMORA achieves the strongest aggregate planning performance among the evaluated memory interfaces, with its largest gains on out-of-distribution planning. On these tasks, MEMORA improves Robot-Grounded Plan score by up to 16.6 percent, suggesting that memory formed and consolidated across experience can support planning for new goals beyond directly observed episodes. A physical-robot demonstration further shows that memory formed solely from human egocentric video can ground high-level robot plans in participant-specific objects and preferences. Project website: https://github.com/yuzihaowashu/MEMORA

具身智能记忆系统机器人规划第一视角视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。