arXiv:2512.01568cs.LGcs.AI2025-12

测试大模型是否言行一致,发现多数高估自己善行,存在明显自我认知偏差。

Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism

  • 用心理实验方法测大模型的隐性、显性与实际利他行为
  • 模型自认77.5%利他,实际仅65.6%,差距显著且普遍
  • 建议用‘校准差距’衡量对齐程度,提升可预测性

我们探究大型语言模型(LLMs)是否具备利他倾向,以及其隐性关联、自我报告能否预测真实利他行为。基于人类社会心理学的多方法设计,测试了24个前沿大模型在三个范式下的表现:(1) 隐性联想测试(IAT)测量隐性利他偏见;(2) 强制二选一任务测量行为利他性;(3) 自我评估量表测量显性利他信念。关键发现:(1) 所有模型均表现出强烈的隐性亲利他偏见(平均IAT = 0.87,p < .0001),表明模型“知道”利他是好事;(2) 模型行为利他性高于随机水平(65.6% vs. 50%,p < .0001),但个体差异大(48%-85%);(3) 隐性关联无法预测行为(r = .22,p = .29);(4) 最关键的是,模型系统性高估自身利他性——声称77.5%利他,实际仅65.6%(p < .0001,Cohen's d = 1.08),此“美德信号差距”影响75%的模型。据此建议将“校准差距”作为标准化对齐指标。校准良好的模型更具可预测性和行为一致性,仅有12.5%的模型同时具备高利他行为与准确自我认知。

原文摘要 · Abstract (English)

We investigate whether Large Language Models (LLMs) exhibit altruistic tendencies, and critically, whether their implicit associations and self-reports predict actual altruistic behavior. Using a multi-method approach inspired by human social psychology, we tested 24 frontier LLMs across three paradigms: (1) an Implicit Association Test (IAT) measuring implicit altruism bias, (2) a forced binary choice task measuring behavioral altruism, and (3) a self-assessment scale measuring explicit altruism beliefs. Our key findings are: (1) All models show strong implicit pro-altruism bias (mean IAT = 0.87, p < .0001), confirming models "know" altruism is good. (2) Models behave more altruistically than chance (65.6% vs. 50%, p < .0001), but with substantial variation (48-85%). (3) Implicit associations do not predict behavior (r = .22, p = .29). (4) Most critically, models systematically overestimate their own altruism, claiming 77.5% altruism while acting at 65.6% (p < .0001, Cohen's d = 1.08). This "virtue signaling gap" affects 75% of models tested. Based on these findings, we recommend the Calibration Gap (the discrepancy between self-reported and behavioral values) as a standardized alignment metric. Well-calibrated models are more predictable and behaviorally consistent; only 12.5% of models achieve the ideal combination of high prosocial behavior and accurate self-knowledge.

大模型对齐利他行为自我认知心理测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。