标注论点中的具体情绪类别,发现大模型更倾向误判负面情绪。
Fearful Falcons and Angry Llamas: Emotion Category Annotations of Arguments by Humans and LLMs
- 人工标注德语论点情绪类别,对比三种提示策略与三类大模型表现
- 自动标注对愤怒和恐惧的召回率高但精确率低,显示明显负面情绪偏见
- 首次系统研究论点中离散情绪类别,为情绪影响分析提供新数据
论点会引发情绪,而情绪的类型和强度都会影响其说服效果,例如改变立场的意愿。尽管已有研究关注论点的情绪强度(二元分类),但缺乏对具体情绪类别(如“愤怒”)的标注数据。为此,我们通过众包方式在德语文本论点语料库中收集主观情绪类别标注,并评估基于大语言模型的自动标注方法。具体比较了三种提示策略(零样本、单样本、思维链)在三个指令微调的大模型(Falcon-7b-instruct、Llama-3.1-8B-instruct、GPT-4o-mini)上的表现。同时考察输出空间定义:二元(是否有情绪)、封闭域(从给定标签集中选情绪)、开放域(自由识别情绪)。结果表明,引入情绪类别能提升情绪强度预测能力,凸显离散情绪标注的重要性。所有模型与提示设置下,对愤怒和恐惧的预测均呈现高召回率但低精确率,说明存在强烈的负面情绪偏见。
原文摘要 · Abstract (English)
Arguments evoke emotions, influencing the effect of the argument itself. Not only the emotional intensity but also the category influence the argument's effects, for instance, the willingness to adapt stances. While binary emotionality has been studied in arguments, there is no work on discrete emotion categories (e.g., "Anger") in such data. To fill this gap, we crowdsource subjective annotations of emotion categories in a German argument corpus and evaluate automatic LLM-based labeling methods. Specifically, we compare three prompting strategies (zero-shot, one-shot, chain-of-thought) on three large instruction-tuned language models (Falcon-7b-instruct, Llama-3.1-8B-instruct, GPT-4o-mini). We further vary the definition of the output space to be binary (is there emotionality in the argument?), closed-domain (which emotion from a given label set is in the argument?), or open-domain (which emotion is in the argument?). We find that emotion categories enhance the prediction of emotionality in arguments, emphasizing the need for discrete emotion annotations in arguments. Across all prompt settings and models, automatic predictions show a high recall but low precision for predicting anger and fear, indicating a strong bias toward negative emotions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。