测试对话中记忆检索的真实表现,发现传统评测遗漏关键问题。
When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

- 设计四种对话式查询测试记忆系统
- 隐含与复合查询存在显著检索差距
- 强检索不等于好回复,适合研究对话智能的学者
大型语言模型日益被用作长时对话代理,推动了记忆系统的研究。然而现有基准主要通过问答式探测评估记忆,而非真实对话场景中的使用。我们引入LOCOMO-CONV,一个基于LoCoMo的对话记忆基准,包含四种查询类型:对话、隐含、反事实和复合。在五个代表性记忆系统上,我们评估了检索召回率和端到端回复质量。实验表明,对话语境暴露了问答基准未发现的显著检索差距,尤其在隐含和复合查询中;多维度查询重写可缩小原始回合记忆的差距,但对抽象记忆无效。此外,强检索并不完全转化为高质量回复,且隐含查询表现出沉默锚定现象——记忆提升了上下文一致性,却未显式呈现正确事实。这些结果指向基于推理的记忆扩展是未来方向,我们发布了辅助性的supportive_memory标注,捕捉原始黄金证据之外的对话有用上下文。
原文摘要 · Abstract (English)
Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory benchmark derived from Lo- CoMo with four query styles: dialog, implicit, counterfactual, and composed. Across five rep- resentative memory systems, we evaluate both retrieval recall and end-to-end response qual- ity. Our experiments show that conversational framing exposes substantial retrieval gaps over- looked by QA benchmarks, especially on im- plicit and composed queries, which multi-facet query rewriting narrows for raw-turn mem- ory but not abstractive memory. We further find that strong retrieval does not fully trans- late into response quality, and that implicit queries exhibit silent grounding, where mem- ory improves contextual grounding without ex- plicitly surfacing the gold fact. These results point to reasoning-based memory elaboration as a promising direction, and we release aux- iliary supportive_memory annotations captur- ing conversationally useful context beyond the original gold evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。