arXiv:2604.20006cs.CL2026-04ACL被引 23

新基准测试揭示个性化代理长期记忆的遗忘问题

From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

论文配图:From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
图 1 · 摘自论文原文
  • 构建跨周至月的对话记忆评估框架,覆盖记事、推理、推荐三任务
  • 发现主流模型频繁使用过时信息,且无法处理知识更新
  • 提出防遗忘评估指标FAMA,适合研究长期记忆的开发者

能长期与用户互动的个性化代理需在会话间保持持续记忆并随情境更新。然而,现有评测多将长期记忆视为历史对话的事实召回,难以反映记忆的长期固化或频繁更新能力。我们提出Memora,一个涵盖数周至数月对话的长期记忆基准。该基准评估三个基于记忆的任务:记忆、推理和推荐。为确保数据质量,采用自动化记忆校验与人工评估。我们进一步引入“防遗忘记忆准确率”(FAMA)指标,在评估中惩罚对过时或失效记忆的依赖。对四种大模型和六种记忆代理的评测显示,存在频繁重用无效记忆及未能调和动态变化记忆的问题。记忆代理仅带来微弱改进,暴露出当前个性化代理长期记忆机制的不足。

原文摘要 · Abstract (English)

Personalized agents that interact with users over long periods must maintain persistent memory across sessions and update it as circumstances change. However, existing benchmarks predominantly frame long-term memory evaluation as fact retrieval from past conversations, providing limited insight into agents' ability to consolidate memory over time or handle frequent knowledge updates. We introduce Memora, a long-term memory benchmark spanning weeks to months long user conversations. The benchmark evaluates three memory-grounded tasks: remembering, reasoning, and recommending. To ensure data quality, we employ automated memory-grounding checks and human evaluation. We further introduce Forgetting-Aware Memory Accuracy (FAMA), a metric that penalizes reliance on obsolete or invalidated memory when evaluating long-term memory. Evaluations of four LLMs and six memory agents reveal frequent reuse of invalid memories and failures to reconcile evolving memories. Memory agents offer marginal improvements, exposing shortcomings in long-term memory for personalized agents.

长期记忆个性化代理评测基准遗忘检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。