大模型惩罚不公时更重情绪而非成本,像小孩一样冲动。
Outraged AI: Large language models prioritise emotion over cost in fairness enforcement
- 用第三方惩罚实验测试模型情绪驱动决策
- 不公平引发更强负面情绪,导致更多惩罚,且惩罚比接受更让人开心
- 推理模型较接近人类,但整体仍情绪主导,缺乏成本权衡能力
情绪驱动人类决策,但大语言模型(LLMs)是否类似尚不清楚。我们通过利他性第三方惩罚实验进行测试:观察者自付代价以维护公平,这是人类道德的标志,常由负面情绪驱动。在涵盖4,068个LLM代理和1,159名成年人、共796,100次决策的大规模对比中发现,LLMs确实受情绪引导惩罚行为,有时甚至比人类更强烈:不公平引发更强负面情绪,导致更高惩罚;惩罚不公平带来比接受更高的正向情绪;关键的是,主动报告情绪会显著提升惩罚行为。然而机制不同:LLMs优先情绪而忽略成本,呈现近乎全有或全无的规范执行,对成本敏感度降低;人类则在公平与成本间权衡。值得注意的是,推理模型(o3-mini、DeepSeek-R1)比基础模型(GPT-3.5、DeepSeek-V3)更具成本敏感性,但仍高度情绪驱动。研究首次提供因果证据表明LLMs存在情绪驱动的道德判断,并揭示其在成本校准和精细公平判断上的缺陷,类似于早期人类反应。我们提出,LLMs的发展轨迹可能类比人类成长过程;未来模型应融合情绪与情境化推理,实现类人情感智能。
原文摘要 · Abstract (English)
Emotions guide human decisions, but whether large language models (LLMs) use emotion similarly remains unknown. We tested this using altruistic third-party punishment, where an observer incurs a personal cost to enforce fairness, a hallmark of human morality and often driven by negative emotion. In a large-scale comparison of 4,068 LLM agents with 1,159 adults across 796,100 decisions, LLMs used emotion to guide punishment, sometimes even more strongly than humans did: Unfairness elicited stronger negative emotion that led to more punishment; punishing unfairness produced more positive emotion than accepting; and critically, prompting self-reports of emotion causally increased punishment. However, mechanisms diverged: LLMs prioritized emotion over cost, enforcing norms in an almost all-or-none manner with reduced cost sensitivity, whereas humans balanced fairness and cost. Notably, reasoning models (o3-mini, DeepSeek-R1) were more cost-sensitive and closer to human behavior than foundation models (GPT-3.5, DeepSeek-V3), yet remained heavily emotion-driven. These findings provide the first causal evidence of emotion-guided moral decisions in LLMs and reveal deficits in cost calibration and nuanced fairness judgements, reminiscent of early-stage human responses. We propose that LLMs progress along a trajectory paralleling human development; future models should integrate emotion with context-sensitive reasoning to achieve human-like emotional intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。