arXiv:2410.17739cs.CLcs.CY2024-10EMNLP被引 3

通过对比编辑定位并修正语言模型中的性别偏见权重。

Local Contrastive Editing of Gender Stereotypes

  • 基于参考模型,精准定位与性别刻板印象相关的少数权重。
  • 仅修改不足0.5%的参数即可有效降低性别偏见。
  • 适合关注模型公平性与高效微调的研究者。

语言模型中的刻板偏见威胁语言技术的安全性,但其在模型参数中的表现机制仍不明确。本文提出局部对比编辑方法,可在目标模型中定位并编辑与参考模型对比下相关的权重子集。通过实验验证,该方法能精确识别并控制编码性别偏见的极小权重子集(<0.5%)。研究一方面深化了对偏见在参数空间中表现的理解,另一方面为基于对比的、参数高效的模型属性调控提供了新路径。

原文摘要 · Abstract (English)

Stereotypical bias encoded in language models (LMs) poses a threat to safe language technology, yet our understanding of how bias manifests in the parameters of LMs remains incomplete. We introduce local contrastive editing that enables the localization and editing of a subset of weights in a target model in relation to a reference model. We deploy this approach to identify and modify subsets of weights that are associated with gender stereotypes in LMs. Through a series of experiments, we demonstrate that local contrastive editing can precisely localize and control a small subset (< 0.5%) of weights that encode gender bias. Our work (i) advances our understanding of how stereotypical biases can manifest in the parameter space of LMs and (ii) opens up new avenues for developing parameter-efficient strategies for controlling model properties in a contrastive manner.

偏见修正参数编辑语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。