首个评估大模型情感理解认知过程的基准,揭示其在情绪推理上的短板。
CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

- 基于评估理论构建包含完整推理链的多视角情感评估框架。
- 模型在正向情绪识别和推理链任务上表现落后于人类。
- 适合关注大模型情感认知能力评估的研究者使用。
情感理解是大模型有效与人类交互的核心能力,但现有评估范式依赖离散情绪标签预测,未能捕捉情绪生成背后的认知过程。基于评估理论,我们提出CAREBench,首个涵盖真实叙事中第一人称与第三人称视角的完整推理链标注基准,包含评估推理、评估评分及多标签情绪标注。我们设计了过程级评估框架,在六种大模型上围绕四个研究问题开展系统实验。结果发现:更强模型在部分任务上可达到或超过人类表现,但在评估推理和正向情绪识别上仍显不足;各模型在推理链不同步骤的表现及对评估干预的敏感性存在差异;当前模型尚未内化捕捉人类主观异质性的机制。这些发现表明,下游情绪预测指标可能高估大模型的真实情感理解能力,CAREBench为更诊断性地评估大模型情感能力提供了基础。
原文摘要 · Abstract (English)
Emotion understanding is a core capability for LLMs to interact effectively with humans, yet existing evaluation paradigms rely on discrete emotion label prediction and fail to capture the cognitive processes underlying emotion generation. Grounded in appraisal theory, we introduce CAREBench, the first benchmark with complete inferential chain annotations from both first- and third-person perspectives on real-world narratives, spanning appraisal reasoning, appraisal ratings, and multi-label emotion annotation. We propose a process-level evaluation framework and conduct systematic experiments across six LLMs organized around four research questions. We find that stronger models match or surpass human observers on certain tasks, yet fall short on appraisal reasoning and positive emotion recognition; performance across chain steps and sensitivity to appraisal interventions exhibit dissociations across models; and current models have not internalized the mechanisms needed to capture human subjective heterogeneity. These findings suggest that downstream emotion prediction metrics may overestimate LLMs' true emotion understanding, and CAREBench provides a foundation for more diagnostically informative evaluation of LLMs' affective cognitive capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。