用小型模型模拟大模型偏见演化,大幅降低去偏研究成本。
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
- 用微型BERT模型在小数据集上模拟偏见形成过程。
- 将预训练成本从500+小时降至30小时以内。
- 适合想快速验证去偏方法的研究者和团队。
预训练语言模型近年来在社会应用和训练成本上均显著增长,但其规模扩张限制了对偏见的理解与缓解。由于重新训练成本过高,现有去偏工作多依赖事后或掩码策略,难以触及偏见根源。本文提出使用低成本代理模型——BabyLMs,即在小而可变语料上训练的紧凑BERT类模型,以模拟大型模型的偏见获取与学习动态。实验表明,尽管体积极小,BabyLMs在内在偏见形成与性能发展上与标准BERT模型高度一致,且在多种模型内与模型后去偏方法中表现出强相关性。利用这一相似性,我们通过BabyLMs开展预训练阶段去偏实验,复现了已有成果,并揭示了性别失衡与毒性内容对偏见形成的显著影响。结果证明,BabyLMs能有效作为大型模型的去偏研究沙盒,将预训练成本从超过500 GPU小时降至不足30 GPU小时,为去偏研究的普及化提供可能,支持更快速、更开放的方法探索。
原文摘要 · Abstract (English)
Pre-trained language models (LMs) have, over the last few years, grown substantially in both societal adoption and training costs. This rapid growth in size has constrained progress in understanding and mitigating their biases. Since re-training LMs is prohibitively expensive, most debiasing work has focused on post-hoc or masking-based strategies, which often fail to address the underlying causes of bias. In this work, we seek to democratise pre-model debiasing research by using low-cost proxy models. Specifically, we investigate BabyLMs, compact BERT-like models trained on small and mutable corpora that can approximate bias acquisition and learning dynamics of larger models. We show that BabyLMs display closely aligned patterns of intrinsic bias formation and performance development compared to standard BERT models, despite their drastically reduced size. Furthermore, correlations between BabyLMs and BERT hold across multiple intra-model and post-model debiasing methods. Leveraging these similarities, we conduct pre-model debiasing experiments with BabyLMs, replicating prior findings and presenting new insights regarding the influence of gender imbalance and toxicity on bias formation. Our results demonstrate that BabyLMs can serve as an effective sandbox for large-scale LMs, reducing pre-training costs from over 500 GPU-hours to under 30 GPU-hours. This provides a way to democratise pre-model debiasing research and enables faster, more accessible exploration of methods for building fairer LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。