arXiv:2506.20020cs.AIcs.CL2025-06ACL被引 15

给大模型分配人格后,会像人一样因身份认同产生偏见推理。

Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning

  • 通过设定政治等8种人格角色,测试模型推理倾向。
  • 带人格的模型辨别虚假新闻能力下降最高达9%。
  • 政治人格使模型更倾向支持符合身份的科学结论,难通过提示修正。

人类推理常受身份保护等动机影响,导致认知偏见,尤其在气候变迁、疫苗安全等议题上加剧社会分裂。已有研究发现大语言模型(LLMs)也存在类似偏见,但其是否会选择性地得出与身份一致的结论仍不明确。本文通过在4类社会人口属性中设定8种人格,测试8个主流模型(开源与专有)在两项人类实验任务中的表现:虚假新闻真伪判断和科学数据评估。结果表明,赋予人格的模型在辨别虚假信息上能力下降最多达9%;政治人格在枪支管控科学证据评估中,当真实结论与身份一致时,正确率提升最高达90%。基于提示的去偏方法基本无效。本研究首次实证显示,人格赋值后的模型表现出难以通过常规提示缓解的人类式动机性推理,引发对模型及人类加剧身份认同偏见的双重担忧。

原文摘要 · Abstract (English)

Reasoning in humans is prone to biases due to underlying motivations like identity protection, that undermine rational decision-making and judgment. This \textit{motivated reasoning} at a collective level can be detrimental to society when debating critical issues such as human-driven climate change or vaccine safety, and can further aggravate political polarization. Prior studies have reported that large language models (LLMs) are also susceptible to human-like cognitive biases, however, the extent to which LLMs selectively reason toward identity-congruent conclusions remains largely unexplored. Here, we investigate whether assigning 8 personas across 4 political and socio-demographic attributes induces motivated reasoning in LLMs. Testing 8 LLMs (open source and proprietary) across two reasoning tasks from human-subject studies -- veracity discernment of misinformation headlines and evaluation of numeric scientific evidence -- we find that persona-assigned LLMs have up to 9% reduced veracity discernment relative to models without personas. Political personas specifically are up to 90% more likely to correctly evaluate scientific evidence on gun control when the ground truth is congruent with their induced political identity. Prompt-based debiasing methods are largely ineffective at mitigating these effects. Taken together, our empirical findings are the first to suggest that persona-assigned LLMs exhibit human-like motivated reasoning that is hard to mitigate through conventional debiasing prompts -- raising concerns of exacerbating identity-congruent reasoning in both LLMs and humans.

大模型认知偏见动机推理人格建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。