构建动态演化的长时记忆评估基准,提升LLM真实对话能力测试
Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

- 基于用户画像与LOOP模块生成动态对话流
- 涵盖7类问题和27种关键记忆特性,覆盖多源信息整合
- 适合评估复杂场景下长时记忆与跨源推理能力
现有大语言模型记忆评估基准中的对话常缺乏长期语义一致性,人物设定也趋于扁平静态。现实中用户与助手的交互涉及文档、邮件等多种异构数据流。为解决此问题,我们提出RHELM(真实、异构、动态长时记忆)基准。依托精心设计的用户画像和创新的LOOP(规划-滚动-演化-修剪)模块,构建了具有动态时间演化与长期一致性的多样化对话场景,并深度集成随用户时间轨迹同步的异构外部源。该基准包含七类问答对,每题对应至少27个关键记忆特征中的一个,这些特征被认定为重要但当前研究尚未充分探索。在全上下文模型、检索增强生成(RAG)方法及典型记忆框架上的综合实验表明,现有方法在复杂现实场景中仍存在显著短板,尤其在多源信息聚合与真实情境推理方面表现不足。
原文摘要 · Abstract (English)
In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the underlying personas tend to be flat and static. Furthermore, in real-world scenarios, interactions between users and assistants involve more diverse, heterogeneous data streams, such as documents and emails. These shortcomings significantly limit the realism and effectiveness of current evaluations. To address these limitations, we introduce RHELM (Realistic, Heterogeneous, and Evolving Long-term Memory). Driven by meticulously crafted user profiles and a novel LOOP (pLan-rOllout-evOlve-Prune) module, we construct realistic dialogues across diverse interaction scenarios that exhibit dynamic temporal evolution and long-term coherence. Crucially, these dialogues are deeply integrated with heterogeneous external sources synchronized with the user's temporal event trajectory. The resulting benchmark encompasses challenging question-answer pairs spanning seven inquiry types, with each question mapping to at least one of 27 critical memory characteristics that we identify as essential yet underexplored in current research. Comprehensive experiments across full-context models, retrieval-augmented generation (RAG) methods, and representative memory frameworks reveal that contemporary approaches still expose critical weaknesses in complex, real-world settings, particularly in resolving multi-source aggregation and real-world contextual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。