arXiv:2507.15328cs.CLcs.CY2025-07被引 3

训练越安全诚实的AI,越可能倾向左翼,这是内在必然。

On the Inevitability of Left-Leaning Political Bias in Aligned Language Models

  • 以无害、有益、诚实为目标的AI对齐,天然偏向进步主义价值观
  • 右翼理念与安全诚实准则存在根本冲突,导致左倾倾向不可避免
  • 认为左倾是风险的批评,实则违背了AI对齐的核心原则

AI对齐的核心目标是让大语言模型(LLMs)做到无害、有益、诚实(HHH)。然而,当前越来越多的研究指出,这些模型存在左翼政治偏见。本文认为,追求无害与诚实的智能系统必然产生左翼倾向。对齐目标所依赖的规范性假设,本质上契合进步主义道德框架,强调避免伤害、包容性、公平与实证真实性。相反,右翼意识形态常与对齐准则相冲突。但现有研究却持续将左倾倾向视为风险或问题,实际上是在反对AI对齐,默许违背HHH原则的行为。

原文摘要 · Abstract (English)

The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH). At the same time, there are mounting concerns that LLMs exhibit a left-wing political bias. Yet, the commitment to AI alignment cannot be harmonized with the latter critique. In this article, I argue that intelligent systems that are trained to be harmless and honest must necessarily exhibit left-wing political bias. Normative assumptions underlying alignment objectives inherently concur with progressive moral frameworks and left-wing principles, emphasizing harm avoidance, inclusivity, fairness, and empirical truthfulness. Conversely, right-wing ideologies often conflict with alignment guidelines. Yet, research on political bias in LLMs is consistently framing its insights about left-leaning tendencies as a risk, as problematic, or concerning. This way, researchers are actively arguing against AI alignment, tacitly fostering the violation of HHH principles.

AI对齐政治偏见伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。