用心理认知理论评估大模型情绪推理能力,突破表面线索局限。
Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models
- 基于认知评估理论构建双向推理数据集,测试模型从上下文推情绪、从情绪反推上下文的能力。
- 模型在情境结果与情绪关联上表现不佳,说明其情绪理解仍停留在表面。
- 为情感推理提供心理学框架,适合关注模型心智理解能力的研究者参考。
情绪识别数据集通常包含明显的表情线索,可直接用于预测文本中的情绪。然而,部分文本含有隐含的上下文线索,蕴含丰富的情感语义,需高阶推理才能推断情绪状态,而非仅依赖表层信息。本研究超越表层感知特征,基于心理理论(ToM)框架,探讨大语言模型如何利用上下文信息推理他人情绪状态。依据认知评估理论,我们构建了一个专门的ToM评估数据集1,用于评估正向推理(从上下文到情绪)与逆向推理(从情绪到推断上下文)。结果显示,尽管模型具备一定推理能力,但在将情境结果与评价要素关联到特定情绪方面表现较差。该工作强调了在情绪推理任务中引入心理学理论对模型训练与评估的重要性。
原文摘要 · Abstract (English)
Datasets used for emotion recognition tasks typically contain overt cues that can be used in predicting the emotions expressed in a text. However, one challenge is that texts sometimes contain covert contextual cues that are rich in affective semantics, which warrant higher-order reasoning abilities to infer emotional states, not simply the emotions conveyed. This study advances beyond surface-level perceptual features to investigate how large language models (LLMs) reason about others' emotional states using contextual information, within a Theory-of-Mind (ToM) framework. Grounded in Cognitive Appraisal Theory, we curate a specialized ToM evaluation dataset1 to assess both forward reasoning - from context to emotion- and backward reasoning - from emotion to inferred context. We showed that LLMs can reason to a certain extent, although they are poor at associating situational outcomes and appraisals with specific emotions. Our work highlights the need for psychological theories in the training and evaluation of LLMs in the context of emotion reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。