用合成数据训练语言模型会悄悄放大偏见,导致公平性崩溃。
The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data
- 在合成数据上反复训练模型,观察偏见演化过程。
- 偏见加剧先于标准指标恶化,最早可被检测到。
- 适合关注模型伦理与数据安全的研究者阅读。
在人工生成数据上训练的语言模型已表现出模型坍塌现象,导致性能显著下降。随着合成内容不断污染语言模型的训练语料,这引发了对开放数据在持续预训练中使用的重大担忧。尽管已有研究揭示了语言模型中的模型坍塌,但其对预训练模型中已有社会偏见的影响尚不明确。由于语言模型会复现并放大性别、种族等刻板印象,基于自生成数据的递归训练可能形成自我强化的反馈循环,使偏见随代际传播而不断增强。我们称此假设现象为公平性坍塌(fairness collapse)。本研究构建了受控训练环境,使用Bias in Bios数据集在合成数据上重复训练模型。实验结果显示,公平性退化在标准语言建模指标出现明显恶化前即已显现。这一发现凸显了合成数据污染语言模型训练的关键风险:偏见可能在无明显信号的情况下悄然增强。
原文摘要 · Abstract (English)
Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation. As synthetic content increasingly contaminates the training corpora of language models, this raises critical concerns about the use of open data in continued pretraining. Although previous work has demonstrated model collapse in language models, it remains unclear whether exposure to synthetic data amplifies or attenuates the social biases already present in pretrained models. Because language models are known to reproduce and amplify demographic stereotypes, recursive training on self-generated data may create a self-reinforcing feedback loop in which biased associations become progressively stronger across generations. We call this hypothesized phenomenon fairness collapse. In this work, we construct controlled training regimes in which models are repeatedly trained on synthetic data using the Bias in Bios dataset. Across experiments, we observe a consistent and concerning pattern: fairness degradation emerges before substantial degradation is reflected by standard language-modeling metrics. This result highlights a critical risk associated with synthetic data contamination in language model training: bias can increase silently before strong indicators of model collapse become apparent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。