arXiv:2602.01885cs.CLcs.AI2026-02中稿 · The Web Conference被引 11

评测大模型在长期情感支持中的记忆能力,发现显式记忆对减少幻觉至关重要。

ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support

  • 构建五维记忆能力评测框架,覆盖信息提取到用户建模
  • 实测显示显式记忆可显著降低幻觉率,提升个性化效果
  • 适合研究长时对话、情感计算与记忆增强系统的开发者

大型语言模型在对话系统中展现出潜力,但在在线情感支持等长期网络服务中仍受限于鲁棒性不足的长期记忆。现有基准多聚焦静态事实检索,难以评估用户信息分散、隐含且持续演变的真实场景。为此,我们提出ES-MemEval,一个涵盖问答、摘要和对话生成任务的综合性评测基准,系统评估五项核心记忆能力:信息提取、时间推理、冲突检测、回避策略与用户建模。同时构建EvoEmo多轮情感支持数据集,捕捉碎片化、隐含的用户披露与动态演化状态。对开源长上下文模型、商用模型及检索增强(RAG)模型的实验表明,显式长期记忆对降低幻觉和实现有效个性化至关重要;而RAG虽提升事实一致性,却在时间动态与状态演化上表现不佳。这些发现揭示了当前范式的潜力与局限,推动更稳健的记忆与检索融合机制发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown strong potential as conversational agents. Yet, their effectiveness remains limited by deficiencies in robust long-term memory, particularly in complex, long-term web-based services such as online emotional support. However, existing long-term dialogue benchmarks primarily focus on static and explicit fact retrieval, failing to evaluate agents in critical scenarios where user information is dispersed, implicit, and continuously evolving. To address this gap, we introduce ES-MemEval, a comprehensive benchmark that systematically evaluates five core memory capabilities: information extraction, temporal reasoning, conflict detection, abstention, and user modeling, in long-term emotional support settings, covering question answering, summarization, and dialogue generation tasks. To support the benchmark, we also propose EvoEmo, a multi-session dataset for personalized long-term emotional support that captures fragmented, implicit user disclosures and evolving user states. Extensive experiments on open-source long-context, commercial, and retrieval-augmented (RAG) LLMs show that explicit long-term memory is essential for reducing hallucinations and enabling effective personalization. At the same time, RAG improves factual consistency but struggles with temporal dynamics and evolving user states. These findings highlight both the potential and limitations of current paradigms and motivate more robust integration of memory and retrieval for long-term personalized dialogue systems.

情感支持长期记忆对话系统评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。