arXiv:2604.07017cs.AI2026-04被引 2

评测模型如何用历史对话记忆理解用户情绪,推动情感智能发展。

A-MBER: Affective Memory Benchmark for Emotion Recognition

论文配图:A-MBER: Affective Memory Benchmark for Emotion Recognition
图 1 · 摘自论文原文
  • 构建多轮对话记忆任务,要求模型基于历史交互推断当前情绪。
  • 在长程隐性情绪、高依赖记忆等场景下表现显著差异,验证记忆关键作用。
  • 适合研究情感计算、对话系统与记忆机制的开发者和研究人员。

能够长期与用户互动的AI助手需理解其当前情绪状态以作出恰当回应,但这一能力尚未充分评估。现有情绪数据集多关注即时情绪,而长期记忆基准则侧重事实回忆或知识更新,难以检验模型是否能利用历史交互来解读当前情绪。为此,我们提出A-MBER——一个面向情绪识别的情感记忆基准,聚焦于基于多会话历史记忆的当前情绪解释。给定一段交互轨迹和指定锚点回合,模型需推断用户当前情绪、识别相关历史证据,并进行有依据的解释。该基准通过分阶段流程构建,包含长周期规划、对话生成、标注、问题设计与最终封装,支持判断、检索与解释任务,并设有模态退化与证据不足等鲁棒性设置。在统一框架下比较局部上下文、长上下文、检索记忆、结构化记忆及黄金证据条件,结果表明A-MBER在设计强调的子集上具有高度区分性,包括长程隐性情绪、高依赖记忆水平、轨迹推理与对抗性设置。研究显示,记忆对情绪解读的作用不仅在于提供更多信息,更在于实现选择性、可落地且情境敏感的记忆使用。

原文摘要 · Abstract (English)

AI assistants that interact with users over time need to interpret the user's current emotional state in order to respond appropriately and personally. However, this capability remains insufficiently evaluated. Existing emotion datasets mainly assess local or instantaneous affect, while long-term memory benchmarks focus largely on factual recall, temporal consistency, or knowledge updating. As a result, current resources provide limited support for testing whether a model can use remembered interaction history to interpret a user's present affective state. We introduce A-MBER, an Affective Memory Benchmark for Emotion Recognition, to evaluate this capability. A-MBER focuses on present affective interpretation grounded in remembered multi-session interaction history. Given an interaction trajectory and a designated anchor turn, a model must infer the user's current affective state, identify historically relevant evidence, and justify its interpretation in a grounded way. The benchmark is constructed through a staged pipeline with explicit intermediate representations, including long-horizon planning, conversation generation, annotation, question construction, and final packaging. It supports judgment, retrieval, and explanation tasks, together with robustness settings such as modality degradation and insufficient-evidence conditions. Experiments compare local-context, long-context, retrieved-memory, structured-memory, and gold-evidence conditions within a unified framework. Results show that A-MBER is especially discriminative on the subsets it is designed to stress, including long-range implicit affect, high-dependency memory levels, trajectory-based reasoning, and adversarial settings. These findings suggest that memory supports affective interpretation not simply by providing more history, but by enabling more selective, grounded, and context-sensitive use of past interaction

情绪识别记忆机制对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。