arXiv:2503.16883cs.CL2025-03

GPT-4可精准标注情绪评估,性能接近甚至超过人类。

Assessing the Reliability and Validity of GPT-4 in Annotating Emotion Appraisal Ratings

  • 用五次生成结果投票提升标注准确率。
  • 模型与人类在21项评估上表现接近,部分优于人类。
  • 长事件描述有助于提高标注精度,适合心理学研究者使用。

情绪评价理论认为情绪源于对事件的主观评估,即评价(appraisal)。该研究考察GPT-4作为阅读者标注21种具体评价维度的能力,在不同提示设置下评估其可靠性和有效性,并与人类标注者对比。结果显示,GPT-4作为阅读者标注器表现优异,性能接近或略超人类;通过五次生成结果的多数投票,其表现可显著提升。此外,单一提示即可有效预测评价分和情绪标签,但增加指令复杂度会降低性能。较长的事件描述能提升模型和人类标注者的准确性。本研究推动了大模型在心理学领域的应用,并提供了优化标注性能的策略。

原文摘要 · Abstract (English)

Appraisal theories suggest that emotions arise from subjective evaluations of events, referred to as appraisals. The taxonomy of appraisals is quite diverse, and they are usually given ratings on a Likert scale to be annotated in an experiencer-annotator or reader-annotator paradigm. This paper studies GPT-4 as a reader-annotator of 21 specific appraisal ratings in different prompt settings, aiming to evaluate and improve its performance compared to human annotators. We found that GPT-4 is an effective reader-annotator that performs close to or even slightly better than human annotators, and its results can be significantly improved by using a majority voting of five completions. GPT-4 also effectively predicts appraisal ratings and emotion labels using a single prompt, but adding instruction complexity results in poorer performance. We also found that longer event descriptions lead to more accurate annotations for both model and human annotator ratings. This work contributes to the growing usage of LLMs in psychology and the strategies for improving GPT-4 performance in annotating appraisals.

情绪识别大模型应用标注优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。