追踪大模型偏见源头,发现训练数据偏见会被放大。
How far can bias go? Tracing bias from pretraining data to alignment
- 通过零样本提示与词元共现分析,追踪偏见传播路径。
- 预训练数据中的性别职业偏见在模型输出中被显著放大。
- 指令微调可缓解部分偏见,但难以根除刻板印象。
随着大语言模型日益应用于用户场景,消除其延续社会不平等的偏见至关重要。尽管已有大量研究聚焦于测量或缓解模型偏见,但对其根源的探究仍不足。本研究以Dolma数据集和OLMo模型为例,分析预训练数据中的性别-职业偏见如何在模型中体现。通过零样本提示与词元共现分析,发现训练数据中的偏见在模型输出中被放大。研究还考察了提示类型、超参数及指令微调对偏见表达的影响:指令微调部分缓解了表征偏见,但仍保留整体性别刻板关联;而超参数与提示变化对偏见影响较小。结果揭示了偏见在整个大模型开发流程中的传递机制,强调应在预训练阶段即着手干预。
原文摘要 · Abstract (English)
As LLMs are increasingly integrated into user-facing applications, addressing biases that perpetuate societal inequalities is crucial. While much work has gone into measuring or mitigating biases in these models, fewer studies have investigated their origins. Therefore, this study examines the correlation between gender-occupation bias in pre-training data and their manifestation in LLMs, focusing on the Dolma dataset and the OLMo model. Using zero-shot prompting and token co-occurrence analyses, we explore how biases in training data influence model outputs. Our findings reveal that biases present in pre-training data are amplified in model outputs. The study also examines the effects of prompt types, hyperparameters, and instruction-tuning on bias expression, finding instruction-tuning partially alleviating representational bias while still maintaining overall stereotypical gender associations, whereas hyperparameters and prompting variation have a lesser effect on bias expression. Our research traces bias throughout the LLM development pipeline and underscores the importance of mitigating bias at the pretraining stage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。