机器人通过用户反馈学会该忘什么,实现长期记忆的智能管理。
Learning to Forget -- Hierarchical Episodic Memory for Lifelong Robot Deployment

- 构建分层记忆结构,用语言模型判断信息重要性并选择性遗忘。
- 真实场景下内存减少45%,查询计算量降低35%,问答准确率不降。
- 越用越准:用户反馈让系统适应个人偏好,第二轮提问准确率提升70%。
机器人在长期部署中需回答用户关于过去经历的问题,如“钥匙放哪了?”或“任务为何失败?”。然而,持续从多模态感知中积累的长期情景记忆(EM)会迅速超出存储容量,使实时查询变得不可行,亟需根据用户对相关性的理解进行选择性遗忘。本文提出 H²-EMV 框架,使人形机器人通过用户交互学习该保留哪些记忆。该方法逐步构建分层情景记忆,基于语言模型和已学习的自然语言规则进行条件化的重要性评估,实现选择性遗忘,并根据用户对被遗忘内容的反馈动态更新规则。在模拟家庭任务和长达20.5小时的真实世界数据(来自 ARMAR-7)上验证表明,H²-EMV 在保持问答准确率的同时,将内存占用减少45%,查询计算量降低35%。关键的是,系统性能随时间提升:通过适应用户特定优先级,第二轮查询的准确率提高了70%,证明所学遗忘机制可实现可扩展、个性化的长期人机协作记忆。
原文摘要 · Abstract (English)
Robots must verbalize their past experiences when users ask "Where did you put my keys?" or "Why did the task fail?" Yet maintaining life-long episodic memory (EM) from continuous multimodal perception quickly exceeds storage limits and makes real-time query impractical, calling for selective forgetting that adapts to users' notions of relevance. We present H$^2$-EMV, a framework enabling humanoids to learn what to remember through user interaction. Our approach incrementally constructs hierarchical EM, selectively forgets using language-model-based relevance estimation conditioned on learned natural-language rules, and updates these rules given user feedback about forgotten details. Evaluations on simulated household tasks and 20.5-hour-long real-world recordings from ARMAR-7 demonstrate that H$^2$-EMV maintains question-answering accuracy while reducing memory size by 45% and query-time compute by 35%. Critically, performance improves over time - accuracy increases 70% in second-round queries by adapting to user-specific priorities - demonstrating that learned forgetting enables scalable, personalized EM for long-term human-robot collaboration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。