分析大模型训练数据的政治倾向,发现偏左内容主导且影响模型立场。
What Is The Political Content in LLMs' Pre- and Post-Training Data?
- 通过抽样与分类分析训练数据政治倾向
- 预训练数据中左倾内容显著多于后训练阶段
- 数据偏见在模型基线阶段就存在且持续传递
大型语言模型生成文本存在政治偏见,但其成因尚不明确。本文从数据角度出发,研究预训练和后训练数据中的政治倾向、数据不平衡、跨数据集相似性及数据-模型对齐问题。通过大规模采样、政治倾向分类和立场检测,发现训练数据系统性偏向左倾,预训练语料比后训练数据包含更多政治相关内容。模型立场与训练数据政治分布高度相关,且不同数据集的预训练内容呈现相似政治分布。政治偏见早在基础模型阶段就已存在,并贯穿后续训练过程。结果表明数据构成是决定模型行为的核心因素,亟需提升训练数据透明度。
原文摘要 · Abstract (English)
Large language models (LLMs) are known to generate politically biased text. Yet, it remains unclear how such biases arise, making it difficult to design effective mitigation strategies. We hypothesize that these biases are rooted in the composition of training data. Taking a data-centric perspective, we formulate research questions on (1) political leaning present in data, (2) data imbalance, (3) cross-dataset similarity, and (4) data-model alignment. We then examine how exposure to political content relates to models' stances on policy issues. We analyze the political content of pre- and post-training datasets of open-source LLMs, combining large-scale sampling, political-leaning classification, and stance detection. We find that training data is systematically skewed toward left-leaning content, with pre-training corpora containing substantially more politically engaged material than post-training data. We further observe a strong correlation between political stances in training data and model behavior, and show that pre-training datasets exhibit similar political distributions despite different curation strategies. In addition, we find that political biases are already present in base models and persist across post-training stages. These findings highlight the central role of data composition in shaping model behavior and motivate the need for greater data transparency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。