提出新评估框架,检测长时记忆系统在冲突下的真实表现。
MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts

- 将记忆有效性视为查询相关的可用性问题,设计三类冲突场景。
- 六种系统在冲突下表现不一,答案正确率常与记忆检索脱节。
- 适合研究长时记忆机制的学者,尤其关注可靠性与鲁棒性。
长时记忆系统使基于大语言模型的对话代理能在多轮会话中保留、检索并应用用户特定信息。然而现有评估主要关注结果性能或时间更新,难以揭示系统在存在冲突选项时如何检索和排序时间有效、事实正确且上下文适用的记忆证据。为此,我们提出 MemConflict,一个诊断框架,将记忆有效性建模为查询依赖的可用性问题。该框架形式化了动态、静态和条件冲突,涵盖时间有效性、事实正确性和上下文适用性。通过结构化用户档案生成可控的长周期历史,引入跨会话冲突,并注入语义相似的干扰项,制造记忆候选间的竞争。由此构建的多会话对话基准支持对最终答案的黑盒评估和对支持记忆检索与排序的白盒分析。六种代表性长时记忆系统的实验显示,不同冲突类型下表现不均,答案正确性常与记忆检索和排序结果分离。敏感性分析表明,更长的历史、干扰项、隐式查询及更大冲突距离均降低性能。诊断发现失败源于缺失支持记忆或检索记忆使用不当。整体上,MemConflict 通过感知检索、认知冲突的方式,推进了长时记忆治理的规范化评估。
原文摘要 · Abstract (English)
Long-term memory systems enable conversational agents based on large language models (LLMs) to retain, retrieve, and apply user-specific information across multi-session interactions. However, existing evaluations mainly assess outcome-level performance or temporal updating, providing limited insight into how systems retrieve and rank temporally valid, factually correct, and contextually applicable memory evidence under conflicting alternatives. To address this gap, we propose MemConflict, a diagnostic framework that treats memory validity as a query-conditioned fitness-for-use problem. MemConflict formalizes dynamic, static, and conditional conflicts over temporal validity, factual correctness, and contextual applicability. It simulates controlled long-horizon histories from structured user profiles, introduces cross-session conflicts, and injects semantically similar distractors to create competition among memory candidates. The resulting multi-session dialogue benchmark supports black-box evaluation of final answers and white-box analysis of supporting-memory retrieval and ranking. Experiments on six representative long-term memory systems show uneven strengths across conflict types, with answer correctness often diverging from memory retrieval and ranking. Sensitivity analyses reveal that longer histories, distractors, implicit queries, and larger conflict distances degrade performance. Diagnostics show failures from missing supporting memories and ineffective use of retrieved memories. Collectively, MemConflict advances principled long-term memory governance through retrieval-aware, conflict-aware reliability assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。