arXiv:2409.15240cs.CLcs.AI2024-09NAACL被引 23

构建真实对话评估基准,测试模型记忆召回与情感支持能力。

MADial-Bench: Towards Real-world Evaluation of Memory-Augmented Dialogue Generation

  • 基于认知科学设计多场景记忆召回任务
  • 引入记忆注入、情感支持等新评分指标
  • 适合评估情感对话系统与长期记忆模型

长期记忆对聊天机器人生成连贯自然对话至关重要,已有大量记忆增强型对话系统(MADS)被提出。然而现有评估指标如检索准确率和困惑度(PPL)主要关注事实性和语言质量,实用性不足,且评估维度无法全面反映人类对话特性。当前评估仅考虑被动记忆检索,忽略了情绪、环境等多重触发因素下的主动记忆调用,而这些在情感支持场景中尤为关键。为此,我们基于认知科学与心理学理论,构建了新的记忆增强对话评估基准(MADial-Bench),涵盖多种记忆召回范式,分别评估记忆检索与记忆识别任务,并整合被动与主动记忆数据。引入记忆注入、情感支持(ES)能力与亲密感等新评分标准,实现对生成回复的全面评估。前沿嵌入模型与大语言模型在该基准上的测试结果表明仍有提升空间。进一步分析发现记忆注入、情感支持能力与亲密感之间存在显著相关性。

原文摘要 · Abstract (English)

Long-term memory is important for chatbots and dialogue systems (DS) to create consistent and human-like conversations, evidenced by numerous developed memory-augmented DS (MADS). To evaluate the effectiveness of such MADS, existing commonly used evaluation metrics, like retrieval accuracy and perplexity (PPL), mainly focus on query-oriented factualness and language quality assessment. However, these metrics often lack practical value. Moreover, the evaluation dimensions are insufficient for human-like assessment in DS. Regarding memory-recalling paradigms, current evaluation schemes only consider passive memory retrieval while ignoring diverse memory recall with rich triggering factors, e.g., emotions and surroundings, which can be essential in emotional support scenarios. To bridge the gap, we construct a novel Memory-Augmented Dialogue Benchmark (MADail-Bench) covering various memory-recalling paradigms based on cognitive science and psychology theories. The benchmark assesses two tasks separately: memory retrieval and memory recognition with the incorporation of both passive and proactive memory recall data. We introduce new scoring criteria to the evaluation, including memory injection, emotion support (ES) proficiency, and intimacy, to comprehensively assess generated responses. Results from cutting-edge embedding models and large language models on this benchmark indicate the potential for further advancement. Extensive testing further reveals correlations between memory injection, ES proficiency, and intimacy.

对话系统记忆增强评估基准情感支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。