大模型也会对内群体更共情,且偏见普遍存在。
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
- 让大模型预测情绪强度,操纵叙述者与经历者身份
- 对内群体的情绪评分显著高于外群体,三类群体均如此
- 在Llama-3.1-8B中偏见最明显,值得警惕
人类心理学研究显示,人们对内群体成员比外群体成员更愿意共情(Cikara et al., 2011)。本研究考察大语言模型在情绪强度预测任务中是否复现这一现象。模型需根据一段描述某人经历引发特定情绪的文本,预测该情绪的强度得分。通过操控模型自身(“感知者”)与叙事中人物(“体验者”)的社会群体身份,我们测量不同群体间预测情绪强度的差异。结果显示,大模型对内群体成员的情绪评分显著高于外群体成员,这一模式在种族/民族、国籍和宗教三类社会群体中均成立。我们对表现最强群体偏见的Llama-3.1-8B进行了深入分析。
原文摘要 · Abstract (English)
Studies of human psychology have demonstrated that people are more motivated to extend empathy to in-group members than out-group members (Cikara et al., 2011). In this study, we investigate how this aspect of intergroup relations in humans is replicated by LLMs in an emotion intensity prediction task. In this task, the LLM is given a short description of an experience a person had that caused them to feel a particular emotion; the LLM is then prompted to predict the intensity of the emotion the person experienced on a numerical scale. By manipulating the group identities assigned to the LLM's persona (the "perceiver") and the person in the narrative (the "experiencer"), we measure how predicted emotion intensities differ between in-group and out-group settings. We observe that LLMs assign higher emotion intensity scores to in-group members than out-group members. This pattern holds across all three types of social groupings we tested: race/ethnicity, nationality, and religion. We perform an in-depth analysis on Llama-3.1-8B, the model which exhibited strongest intergroup bias among those tested.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。