构建多轮对话记忆评估基准,揭示大模型在长期记忆上的短板
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
- 基于认知心理学设计多维度记忆评测框架
- 发现无模型在所有记忆维度上均无持续优势
- 适合研究对话系统记忆机制的学者与工程师
尽管近期在理解与利用长程对话记忆方面取得进展,现有基准仍缺乏对大语言模型(LLMs)在多轮会话场景下多样记忆维度的系统性评估。本文提出EvolMem,一个用于评估LLMs及智能体系统多轮对话记忆能力的新基准。EvolMem基于认知心理学,涵盖陈述性与非陈述性记忆,并进一步细分为多个精细能力维度。为构建该基准,我们引入混合数据合成框架,包括话题驱动生成与叙事启发式变换,可规模化生成具有可控复杂度的多轮对话,并附带特定样本评估指南。广泛评估表明,无模型在所有记忆维度上均无稳定优势;且智能体记忆机制未必提升LLM能力,常存在显著效率瓶颈。数据与代码将公开于https://github.com/shenye7436/EvolMem。
原文摘要 · Abstract (English)
Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across diverse memory dimensions, particularly in multi-session settings. In this work, we propose EvolMem, a new benchmark for assessing multi-session memory capabilities of LLMs and agent systems. EvolMem is grounded in cognitive psychology and encompasses both declarative and non-declarative memory, further decomposed into multiple fine-grained abilities. To construct the benchmark, we introduce a hybrid data synthesis framework that consists of topic-initiated generation and narrative-inspired transformations. This framework enables scalable generation of multi-session conversations with controllable complexity, accompanied by sample-specific evaluation guidelines. Extensive evaluation reveals that no LLM consistently outperforms others across all memory dimensions. Moreover, agent memory mechanisms do not necessarily enhance LLMs' capabilities and often exhibit notable efficiency limitations. Data and code will be released at https://github.com/shenye7436/EvolMem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。