arXiv:2409.05283cs.CLcs.AI2024-09EMNLP被引 48

训练越追求真话,模型越偏向左翼,引发对数据集和偏见的反思。

On the Relationship between Truth and Political Bias in Language Models

  • 用真话数据集训练奖励模型,发现其政治倾向向左。
  • 大模型的左倾偏差更明显,现有开源模型也存在类似问题。
  • 适合关注语言模型对政治立场影响的研究者阅读。

语言模型对齐研究常试图确保模型既有益、无害,又真实且无偏见。然而,同时优化这些目标可能掩盖改进一个方面时对其他方面的影响。本文聚焦于分析语言模型对齐与政治科学中的两个核心概念——真实性和政治偏见之间的关系。我们基于多个流行的真实度数据集训练奖励模型,并评估其政治偏见。结果表明,优化真实性的奖励模型倾向于产生左倾政治偏见。此外,现有开源奖励模型(基于标准人类偏好数据集训练)也表现出类似偏见,且模型越大,偏见越强。这一发现引发了对真实度数据集代表性的质疑,以及在实现真实与政治中立双重目标时的潜在局限性,揭示了语言模型对真理与政治关系的理解机制。

原文摘要 · Abstract (English)

Language model alignment research often attempts to ensure that models are not only helpful and harmless, but also truthful and unbiased. However, optimizing these objectives simultaneously can obscure how improving one aspect might impact the others. In this work, we focus on analyzing the relationship between two concepts essential in both language model alignment and political science: truthfulness and political bias. We train reward models on various popular truthfulness datasets and subsequently evaluate their political bias. Our findings reveal that optimizing reward models for truthfulness on these datasets tends to result in a left-leaning political bias. We also find that existing open-source reward models (i.e., those trained on standard human preference datasets) already show a similar bias and that the bias is larger for larger models. These results raise important questions about the datasets used to represent truthfulness, potential limitations of aligning models to be both truthful and politically unbiased, and what language models capture about the relationship between truth and politics.

语言模型偏见分析真实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。