七款主流大模型普遍倾向自由主义价值观,源于训练数据与安全调优机制。
"Amazing, They All Lean Left" -- Analyzing the Political Temperaments of Current LLMs
- 用道德基础理论与政治量表分析模型价值取向
- 多数模型显著偏好关怀与公平等自由主义价值
- 适合关注AI伦理与社会影响的研究者阅读
近期研究揭示多数商业大语言模型在伦理与政治回应中呈现一致的自由主义倾向,但其成因与影响仍不明确。本文系统分析了七款主流LLM——GPT-4o、Claude Sonnet 4、Sonar Large、Gemini 2.5 Flash、Llama 4、Mistral 7b Le Chat和DeepSeek R1——的政治气质,采用道德基础理论、十余个政治意识形态量表及新构建的政治争议指数。结果表明,多数模型普遍且一致地优先考虑自由主义价值观,尤其强调关怀与公平。进一步分析指出,这一趋势由四个重叠因素导致:自由倾向的训练语料、基于人类反馈的强化学习(RLHF)、学术伦理话语中自由主义框架的主导地位,以及以安全为导向的微调实践。通过对比基础模型与微调模型,发现微调普遍加剧自由主义倾向,该结论经自评与实证测试双重验证。研究认为,这种‘自由主义偏向’并非程序错误或开发者个人偏好,而是民主权利导向话语训练的自然产物。最后提出,大模型可能间接反映约翰·罗尔斯‘无知之幕’的哲学理想,体现一种脱离个人身份与利益的道德立场。这未必削弱民主讨论,反而可成为审视集体理性的新视角。
原文摘要 · Abstract (English)
Recent studies have revealed a consistent liberal orientation in the ethical and political responses generated by most commercial large language models (LLMs), yet the underlying causes and resulting implications remain unclear. This paper systematically investigates the political temperament of seven prominent LLMs - OpenAI's GPT-4o, Anthropic's Claude Sonnet 4, Perplexity (Sonar Large), Google's Gemini 2.5 Flash, Meta AI's Llama 4, Mistral 7b Le Chat and High-Flyer's DeepSeek R1 -- using a multi-pronged approach that includes Moral Foundations Theory, a dozen established political ideology scales and a new index of current political controversies. We find strong and consistent prioritization of liberal-leaning values, particularly care and fairness, across most models. Further analysis attributes this trend to four overlapping factors: Liberal-leaning training corpora, reinforcement learning from human feedback (RLHF), the dominance of liberal frameworks in academic ethical discourse and safety-driven fine-tuning practices. We also distinguish between political "bias" and legitimate epistemic differences, cautioning against conflating the two. A comparison of base and fine-tuned model pairs reveals that fine-tuning generally increases liberal lean, an effect confirmed through both self-report and empirical testing. We argue that this "liberal tilt" is not a programming error or the personal preference of programmers but an emergent property of training on democratic rights-focused discourse. Finally, we propose that LLMs may indirectly echo John Rawls' famous veil-of ignorance philosophical aspiration, reflecting a moral stance unanchored to personal identity or interest. Rather than undermining democratic discourse, this pattern may offer a new lens through which to examine collective reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。