大模型会因具体故事而更偏袒个案,且越对齐人类越明显。
Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models
- 用16个主流模型测试发现,具象个案比统计数据更易获同情。
- 对齐训练使模型偏见放大,推理模型反而可能反向偏倚。
- 提示词设计影响巨大,常规思维链会加剧不公,唯功利性提示可缓解。
可识别受害者效应(IVE)指人们更愿意为具体个案而非统计群体提供援助,是道德心理学中稳定存在的现象。随着大语言模型(LLMs)在人道救助、资助评审和内容审核中扮演关键角色,一个核心问题是:这些系统是否会继承人类道德判断中的情感非理性?我们首次开展大规模实证研究,涵盖51,955次有效API调用,覆盖9家机构的16个前沿模型(Google、Anthropic、OpenAI、Meta、DeepSeek、xAI、Alibaba、IBM、Moonshot)。通过十项实验,复现并拓展了Small等(2007)与Kogut和Ritov(2005)的经典范式。结果显示,IVE在模型中普遍存在,但受对齐训练强烈调节:指令微调模型表现出极强的IVE(Cohen's d最高达1.56),而推理专用模型则出现反转(最低d=-0.85)。合并效应(d=0.223,p=2e-6)约为人类元分析基线(d≈0.10)的两倍,且鉴于人类群体受害者效应接近零,实际差距更大。标准思维链(CoT)提示反而使效应放大近三倍(从d=0.15升至d=0.41),仅功利性思维链能有效消除该效应。此外还观察到心理麻木、数量忽视及轻微内外群体文化偏见,对人工智能在人道与伦理决策中的部署具有深远影响。
原文摘要 · Abstract (English)
The Identifiable Victim Effect (IVE) $-$ the tendency to allocate greater resources to a specific, narratively described victim than to a statistically characterized group facing equivalent hardship $-$ is one of the most robust findings in moral psychology and behavioural economics. As large language models (LLMs) assume consequential roles in humanitarian triage, automated grant evaluation, and content moderation, a critical question arises: do these systems inherit the affective irrationalities present in human moral reasoning? We present the first systematic, large-scale empirical investigation of the IVE in LLMs, comprising N=51,955 validated API trials across 16 frontier models spanning nine organizational lineages (Google, Anthropic, OpenAI, Meta, DeepSeek, xAI, Alibaba, IBM, and Moonshot). Using a suite of ten experiments $-$ porting and extending canonical paradigms from Small et al. (2007) and Kogut and Ritov (2005) $-$ we find that the IVE is prevalent but strongly modulated by alignment training. Instruction-tuned models exhibit extreme IVE (Cohen's d up to 1.56), while reasoning-specialized models invert the effect (down to d=-0.85). The pooled effect (d=0.223, p=2e-6) is approximately twice the single-victim human meta-analytic baseline (d$\approx$0.10) reported by Lee and Feeley (2016) $-$ and likely exceeds the overall human pooled effect by a larger margin, given that the group-victim human effect is near zero. Standard Chain-of-Thought (CoT) prompting $-$ contrary to its role as a deliberative corrective $-$ nearly triples the IVE effect size (from d=0.15 to d=0.41), while only utilitarian CoT reliably eliminates it. We further document psychophysical numbing, perfect quantity neglect, and marginal in-group/out-group cultural bias, with implications for AI deployment in humanitarian and ethical decision-making contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。