arXiv:2410.15234cs.AI2024-10被引 20

LLM在生成文本时会逐步放大政治偏见,且与模型退化机制不同。

Bias Amplification: Large Language Models as Increasingly Biased Media

  • 构建长上下文政治偏见评估基准,通过新闻续写任务检测偏见变化
  • GPT-2在迭代生成中出现明显右倾偏见强化,即使控制模型退化仍持续存在
  • 发现偏见放大与模型退化由不同神经元驱动,提示需针对性干预

模型坍塌(model collapse)指因在合成数据上迭代训练导致性能下降的现象,已有广泛研究。然而,其对大语言模型(LLMs)中偏见放大——即预存社会偏见随生成过程不断加剧——的影响仍严重被忽视,尽管LLMs正日益塑造网络舆论。本文提出一个开放的、生成式且长上下文的基准,专门用于衡量LLMs中的政治偏见放大,基于美国政治新闻数据集设计句子续写任务。对GPT-2的实证研究表明,经过多轮合成训练后,政治偏见显著且一致地强化(如右倾倾向),且这一现象独立于模型坍塌,即使有效控制后者也依然存在。我们评估了过拟合、保留与累积三种缓解策略,均未能根除偏见放大。进一步提出一种机制分析方法,通过回归与统计检验识别推理过程中与特定现象相关的神经元。结果显示,驱动偏见放大的神经元群与驱动模型坍塌的神经元群基本不重叠,表明二者具有根本不同的作用机制。最后,我们补充理论直觉解释两类现象的独立起源,为偏见治理提供精准方向。

原文摘要 · Abstract (English)

Model collapse, a phenomenon characterized by performance degradation due to iterative training on synthetic data, has been widely studied. However, its implications for bias amplification, the progressive intensification of pre-existing societal biases in Large Language Models (LLMs), remain significantly underexplored, despite the growing influence of LLMs in shaping online discourse. In this paper, we introduce a open, generational, and long-context benchmark specifically designed to measure political bias amplification in LLMs, leveraging sentence continuation tasks derived from a comprehensive dataset of U.S. political news. Our empirical study using GPT-2 reveals consistent and substantial political bias intensification (e.g., right-leaning amplification) over iterative synthetic training cycles. We evaluate three mitigation strategies, Overfitting, Preservation, and Accumulation, and demonstrate that bias amplification persists independently of model collapse, even when the latter is effectively controlled. Furthermore, we propose a mechanistic analysis approach that identifies neurons correlated with specific phenomena during inference through regression and statistical tests. This analysis uncovers largely distinct neuron populations driving bias amplification and model collapse, underscoring fundamentally different underlying mechanisms. Finally, we supplement our empirical findings with theoretical intuition that explains the separate origins of these phenomena, guiding targeted strategies for bias mitigation.

偏见放大大模型机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。