arXiv:2602.23944cs.CL2026-02被引 1

评测大模型记忆系统处理情感信息的能力,发现现有系统表现均不理想。

MemEmo: Evaluating Emotion in Memory Systems of Agents

  • 构建人类级情感记忆数据集HLME,从三方面评估记忆系统。
  • 三种任务中无一系统表现稳健,情感记忆处理能力普遍薄弱。
  • 适合关注大模型认知缺陷与情感智能的研究者参考。

记忆系统旨在解决大型语言模型在长时间交互中面临的上下文丢失问题。然而,相较于人类认知,这些系统在处理情感相关信息方面的有效性尚不明确。为弥补这一差距,我们提出了一个增强情感的内存评估基准,用于评估主流及前沿记忆系统的效能在处理情感信息时的表现。我们构建了人类级记忆情感(HLME)数据集,该数据集从三个维度评估记忆系统:情感信息提取、情感记忆更新和情感记忆问答。实验结果表明,在所有三项任务中,没有一个被评估的系统均表现出稳健性能。研究结果为当前记忆系统在处理情感记忆方面的不足提供了客观视角,并为未来的研究方向和系统优化指明了新路径。

原文摘要 · Abstract (English)

Memory systems address the challenge of context loss in Large Language Model during prolonged interactions. However, compared to human cognition, the efficacy of these systems in processing emotion-related information remains inconclusive. To address this gap, we propose an emotion-enhanced memory evaluation benchmark to assess the performance of mainstream and state-of-the-art memory systems in handling affective information. We developed the \textbf{H}uman-\textbf{L}ike \textbf{M}emory \textbf{E}motion (\textbf{HLME}) dataset, which evaluates memory systems across three dimensions: emotional information extraction, emotional memory updating, and emotional memory question answering. Experimental results indicate that none of the evaluated systems achieve robust performance across all three tasks. Our findings provide an objective perspective on the current deficiencies of memory systems in processing emotional memories and suggest a new trajectory for future research and system optimization.

记忆系统情感智能评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。