分析预训练模型中性别偏见的编码模式,发现去偏技术可能加剧内部偏见。
Gender Encoding Patterns in Pretrained Language Model Representations
- 用信息论方法研究多种编码架构如何内化性别偏见
- 去偏技术常使内部表示偏见反而上升,与输出结果矛盾
- 为构建更公平的语言模型提供关键实证依据
预训练语言模型中的性别偏见带来重大社会与伦理挑战。尽管关注度提升,但对不同模型如何内部表征和传播此类偏见仍缺乏系统研究。本研究采用信息论方法,分析多种基于编码器的架构中性别偏见的编码方式。重点考察三个方面:模型如何编码性别信息与偏见;去偏技术与微调对编码偏见的影响及其有效性;模型设计差异如何影响偏见编码。通过严谨系统的调查,发现各类模型存在一致的性别编码模式。令人意外的是,去偏技术往往效果有限,有时反而在内部表示中加剧偏见,尽管输出分布的偏见降低。这揭示了输出偏见缓解与内部表示偏见之间存在脱节。本工作为推进去偏策略、发展更公平的语言模型提供了重要指导。
原文摘要 · Abstract (English)
Gender bias in pretrained language models (PLMs) poses significant social and ethical challenges. Despite growing awareness, there is a lack of comprehensive investigation into how different models internally represent and propagate such biases. This study adopts an information-theoretic approach to analyze how gender biases are encoded within various encoder-based architectures. We focus on three key aspects: identifying how models encode gender information and biases, examining the impact of bias mitigation techniques and fine-tuning on the encoded biases and their effectiveness, and exploring how model design differences influence the encoding of biases. Through rigorous and systematic investigation, our findings reveal a consistent pattern of gender encoding across diverse models. Surprisingly, debiasing techniques often exhibit limited efficacy, sometimes inadvertently increasing the encoded bias in internal representations while reducing bias in model output distributions. This highlights a disconnect between mitigating bias in output distributions and addressing its internal representations. This work provides valuable guidance for advancing bias mitigation strategies and fostering the development of more equitable language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。