评测大模型在隐性约束下的长期记忆能力,发现现有方法有遗漏。
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
- 设计新基准,评估模型在提示与语义脱节时的隐性约束记忆。
- 多模型测试显示认知记忆仍严重不足,传统评测方法无法发现缺陷。
- 适合研究长对话、具身智能和可信AI的学者使用。
基于大模型的对话系统需要长期对话记忆能力,但现有评测主要关注表面事实回忆。在真实交互中,恰当回应常依赖未明说的用户状态、目标或价值观等隐性约束,这些约束不会被直接提问。为此,我们提出 extbf{LoCoMo-Plus} 基准,用于评估在提示-触发语义脱节场景下模型对潜在约束的保留与应用能力。实验表明,传统字符串匹配指标和显式任务提示与该场景不匹配,我们提出基于约束一致性的统一评测框架。在多种骨干模型、检索方法和记忆系统上的测试显示,认知记忆仍具挑战性,且现有基准未能捕捉到此类失败。代码与评测框架已开源:https://github.com/xjtuleeyf/Locomo-Plus。
原文摘要 · Abstract (English)
Long-term conversational memory is a core capability for LLM-based dialogue systems, yet existing benchmarks and evaluation protocols primarily focus on surface-level factual recall. In realistic interactions, appropriate responses often depend on implicit constraints such as user state, goals, or values that are not explicitly queried later. To evaluate this setting, we introduce \textbf{LoCoMo-Plus}, a benchmark for assessing cognitive memory under cue--trigger semantic disconnect, where models must retain and apply latent constraints across long conversational contexts. We further show that conventional string-matching metrics and explicit task-type prompting are misaligned with such scenarios, and propose a unified evaluation framework based on constraint consistency. Experiments across diverse backbone models, retrieval-based methods, and memory systems demonstrate that cognitive memory remains challenging and reveals failures not captured by existing benchmarks. Our code and evaluation framework are publicly available at: https://github.com/xjtuleeyf/Locomo-Plus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。