测试大模型对细粒度情绪的理解能力,发现其与人类真实表达仍有差距。
Fluent but Unfeeling: The Emotional Blind Spots of Language Models
- 构建251种自述情绪标签的Reddit数据集EXPRESS,支持细粒度评估
- 多数大模型在预测具体情绪时与人类自我披露不一致,准确率不足
- 适合关注情绪理解、心理计算和人机共情的研究者参考
大型语言模型(LLMs)在自然语言理解中的多功能性使其在心理健康研究中日益流行。尽管许多研究探讨了LLMs在情绪识别方面的能力,但一个关键空白仍在于评估它们是否能在细粒度层面与人类情感保持一致。现有研究通常将情绪分类为预定义的有限类别,忽视了更细微的情感表达。为此,我们引入EXPRESS,一个从Reddit社区收集的基准数据集,包含251种细粒度的自述情绪标签。我们的综合评估框架分析预测的情绪词,并基于成熟情绪理论将其分解为八种基本情绪,实现细粒度对比。在多种提示设置下对主流LLMs进行系统测试发现,准确预测与人类自述相符的情绪仍具挑战性。定性分析进一步显示,虽然某些LLMs生成的情绪术语符合既定情绪理论和定义,但在捕捉上下文线索方面,仍不如人类自述有效。这些发现揭示了LLMs在细粒度情绪对齐上的局限性,为未来提升其上下文理解能力提供了洞见。
原文摘要 · Abstract (English)
The versatility of Large Language Models (LLMs) in natural language understanding has made them increasingly popular in mental health research. While many studies explore LLMs' capabilities in emotion recognition, a critical gap remains in evaluating whether LLMs align with human emotions at a fine-grained level. Existing research typically focuses on classifying emotions into predefined, limited categories, overlooking more nuanced expressions. To address this gap, we introduce EXPRESS, a benchmark dataset curated from Reddit communities featuring 251 fine-grained, self-disclosed emotion labels. Our comprehensive evaluation framework examines predicted emotion terms and decomposes them into eight basic emotions using established emotion theories, enabling a fine-grained comparison. Systematic testing of prevalent LLMs under various prompt settings reveals that accurately predicting emotions that align with human self-disclosed emotions remains challenging. Qualitative analysis further shows that while certain LLMs generate emotion terms consistent with established emotion theories and definitions, they sometimes fail to capture contextual cues as effectively as human self-disclosures. These findings highlight the limitations of LLMs in fine-grained emotion alignment and offer insights for future research aimed at enhancing their contextual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。