首个情绪陪伴对话系统评测基准,可量化模型情感支持能力差异
MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
- 构建四层架构评估框架,定义情绪陪伴对话系统标准
- 30个主流模型测试显示现有系统缺乏深度情感支持能力
- 适合对话系统开发者、人机交互研究者用于优化用户体验
随着大语言模型的快速发展,对话系统正从信息工具转向情感陪伴角色,迎来情绪陪伴对话系统(ECDs)时代,为用户提供个性化情感支持。然而,该领域尚无明确界定与系统的评估标准。为此,我们首先提出ECDs的正式定义。基于此理论及“能力层-任务层(三层)-数据层-方法层”设计原则,构建首个ECD评估基准——MoodBench 1.0。通过对30个主流模型的广泛评估,验证了MoodBench 1.0具有优异的区分效度,能有效量化模型间的情感陪伴能力差异。结果揭示当前模型在深层情感陪伴方面存在明显不足,为未来技术优化提供方向,显著助力开发者提升ECDs用户体验。
原文摘要 · Abstract (English)
With the rapid development of Large Language Models, dialogue systems are shifting from information tools to emotional companions, heralding the era of Emotional Companionship Dialogue Systems (ECDs) that provide personalized emotional support for users. However, the field lacks clear definitions and systematic evaluation standards for ECDs. To address this, we first propose a definition of ECDs with formal descriptions. Then, based on this theory and the design principle of "Ability Layer-Task Layer (three level)-Data Layer-Method Layer", we design and implement the first ECD evaluation benchmark - MoodBench 1.0. Through extensive evaluations of 30 mainstream models, we demonstrate that MoodBench 1.0 has excellent discriminant validity and can effectively quantify the differences in emotional companionship abilities among models. Furthermore, the results reveal current models' shortcomings in deep emotional companionship, guiding future technological optimization and significantly aiding developers in enhancing ECDs' user experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。