LLM评价道德行为时表面像人,实则动机判断不同。
Human-like moral judgments conceal divergent motive attributions in large language models

- 用同一套问题测试人类和大模型的道德判断
- 模型认为举报者更利他、更无私,动机关联性弱
- 适合用于心理研究模拟时需检验深层认知模式
大型语言模型(LLMs)被用于心理学研究中模拟人类参与者。我们考察了这些模型在复制人类对举报者道德评价的同时,是否也复制了伴随其后的动机归因。五种LLM及两组人类样本(分别为N = 125 和 N = 742)评估了一名医生:要么对欺诈性账单保持沉默,要么向医院、监管机构或报纸举报。模型复现了人类对医生道德品质的排名,但将举报者视为更助人、更少自利、更少敌意。在五种模型中,有四种显示竞争性动机与道德评价的关联性较弱。当提示中重现两组人类样本的故事叙述和人口统计特征时,模型评分变化极小,但此比较无法排除视角效应。因此,平均评分的一致性可能掩盖动机归因、判断间关系及情境敏感性的差异。验证LLM作为模拟参与者,需检测心理上有意义的响应模式,而非仅依赖平均一致性。
原文摘要 · Abstract (English)
Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physician who either remained silent about fraudulent billing or reported it to a hospital, regulator, or newspaper. Models reproduced the human ranking of the physician's moral character but portrayed whistleblowers as more helpful, less self-interested, and less hostile. In four of five models, competitive motives were less strongly associated with moral-character judgments. Model ratings changed little when prompts reproduced the narratives and demographic profiles of both human samples, although this comparison cannot isolate a perspective effect. Thus, agreement in average ratings can conceal differences in attributed motives, relationships among judgments, and sensitivity to context. Validating LLMs as simulated participants therefore requires testing psychologically informative response patterns, not average agreement alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。