arXiv:2409.13359cs.CLcs.AI2024-09中稿 · ACL

构建新基准评估大模型共情能力,涵盖四类情感任务

EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models

  • 设计四项情感任务:关键事件、混合事件、隐含情绪与意图识别
  • 发现大模型在隐含情绪识别上表现较差,共情回应能力有限
  • 适合研究情感智能、人机共情的学者与开发者参考

大型语言模型(LLMs)的情感智能在自然语言处理中至关重要。然而,以往研究多集中于基础的情绪分析任务(如情绪识别),难以全面评估模型的整体情感智能水平。为此,本文提出名为 EmotionQueen 的新型评估框架,包含四项独特任务:关键事件识别、混合事件识别、隐含情绪识别和意图识别。模型需识别重要事件或隐含情绪,并生成具有共情力的回应。我们还设计了两项指标,用于评估模型在情绪相关语句上的识别与回应能力。实验揭示了大模型在情感智能方面的显著能力与局限性。

原文摘要 · Abstract (English)

Emotional intelligence in large language models (LLMs) is of great importance in Natural Language Processing. However, the previous research mainly focus on basic sentiment analysis tasks, such as emotion recognition, which is not enough to evaluate LLMs' overall emotional intelligence. Therefore, this paper presents a novel framework named EmotionQueen for evaluating the emotional intelligence of LLMs. The framework includes four distinctive tasks: Key Event Recognition, Mixed Event Recognition, Implicit Emotional Recognition, and Intention Recognition. LLMs are requested to recognize important event or implicit emotions and generate empathetic response. We also design two metrics to evaluate LLMs' capabilities in recognition and response for emotion-related statements. Experiments yield significant conclusions about LLMs' capabilities and limitations in emotion intelligence.

情感智能共情评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。