arXiv:2604.19125cs.CL2026-04ACL被引 1

情绪会显著影响大模型的道德判断,且效果可反转结论。

Do Emotions Influence Moral Judgment in Large Language Models?

论文配图:Do Emotions Influence Moral Judgment in Large Language Models?
图 1 · 摘自论文原文
  • 通过情绪诱导使模型在道德情境中产生情感反应
  • 正情绪提升可接受性,负情绪降低可接受性,20%判断被反转
  • 高能力模型更不易受情绪影响,且存在违背直觉的情绪效应

大语言模型在情绪识别和道德推理方面已得到广泛研究,但情绪对道德判断的影响仍不明确。本文构建了一套情绪诱导流程,将情绪注入道德情境,并在多个数据集和大模型上评估道德可接受性的变化。结果呈现明显方向性:正情绪提升道德可接受性,负情绪则降低,影响足以在最多20%的案例中逆转二元道德判断;模型能力越强,受情绪影响越小。进一步分析发现,某些情绪表现与情感极性预测相反(如悔恨反而提高可接受性)。人类标注对照研究显示,人类无此类系统性偏差,表明当前大模型存在认知对齐差距。

原文摘要 · Abstract (English)

Large language models have been extensively studied for emotion recognition and moral reasoning as distinct capabilities, yet the extent to which emotions influence moral judgment remains underexplored. In this work, we develop an emotion-induction pipeline that infuses emotion into moral situations and evaluate shifts in moral acceptability across multiple datasets and LLMs. We observe a directional pattern: positive emotions increase moral acceptability and negative emotions decrease it, with effects strong enough to reverse binary moral judgments in up to 20% of cases, and with susceptibility scaling inversely with model capability. Our analysis further reveals that specific emotions can sometimes behave contrary to what their valence would predict (e.g., remorse paradoxically increases acceptability). A complementary human annotation study shows humans do not exhibit these systematic shifts, indicating an alignment gap in current LLMs.

道德推理情绪影响大模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。