发现大模型在中文提示下更倾向支持威权人物,可能强化全球政治偏见。
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
- 用心理量表与领导人偏好评分结合角色模型探测新政治维度
- 中文提示下模型对威权领袖好感度显著上升,且常引用其为榜样
- 揭示模型隐含的政治意识形态,适合关注AI伦理的研究者参考
随着大型语言模型(LLMs)日益融入日常生活与信息生态,其隐含偏见问题持续引发关注。现有研究多聚焦社会人口特征及左右政治光谱,却忽视了模型对民主—威权价值体系的倾向性。本文提出一种新方法,融合F量表(测量威权倾向)、FavScore(评估模型对世界领导人的偏好)以及角色模型探针(分析模型引述的榜样人物)。结果发现,尽管多数模型倾向于民主价值观与领导人,但在中文提示下对威权人物的偏好明显增强;此外,即使在非政治语境中,模型也常将威权人物列为角色榜样。这些发现揭示了大模型可能反映并加剧全球政治意识形态,凸显了超越传统政治轴线评估偏见的重要性。代码已开源:https://github.com/irenestrauss/Democratic-Authoritarian-Bias-LLMs。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become increasingly integrated into everyday life and information ecosystems, concerns about their implicit biases continue to persist. While prior work has primarily examined socio-demographic and left--right political dimensions, little attention has been paid to how LLMs align with broader geopolitical value systems, particularly the democracy--authoritarianism spectrum. In this paper, we propose a novel methodology to assess such alignment, combining (1) the F-scale, a psychometric tool for measuring authoritarian tendencies, (2) FavScore, a newly introduced metric for evaluating model favorability toward world leaders, and (3) role-model probing to assess which figures are cited as general role-models by LLMs. We find that LLMs generally favor democratic values and leaders, but exhibit increased favorability toward authoritarian figures when prompted in Mandarin. Further, models are found to often cite authoritarian figures as role models, even outside explicit political contexts. These results shed light on ways LLMs may reflect and potentially reinforce global political ideologies, highlighting the importance of evaluating bias beyond conventional socio-political axes. Our code is available at: https://github.com/irenestrauss/Democratic-Authoritarian-Bias-LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。