用轻量编辑网络精准消除语言模型偏见,不损伤原有能力
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
- 通过编辑网络局部修改模型参数,实现精准去偏
- 在StereoSet和Crows-Pairs上显著降低偏见,性能几乎不变
- 适合需要快速去偏且保留模型原功能的研究与应用
已有研究证实语言模型存在刻板偏见。现有去偏方法如对抗数据重训练、表示投影和提示工程等,往往无法高效消除偏见或直接修改模型内部偏见表征。为此,我们提出BiasEdit,一种高效的模型编辑方法,通过轻量级编辑网络生成参数更新,以移除语言模型中的刻板偏见。BiasEdit采用去偏损失引导编辑网络对模型部分参数进行局部修改,同时通过保留损失确保编辑过程中保持语言建模能力。在StereoSet和Crows-Pairs上的实验表明,相比其他去偏基线方法,BiasEdit在消除偏见方面更有效、更高效且更具鲁棒性,且对模型通用能力影响极小。此外,我们进行了偏见溯源分析,探究不同模块的偏见分布,并研究了编辑对模型各组件的影响。
原文摘要 · Abstract (English)
Previous studies have established that language models manifest stereotyped biases. Existing debiasing strategies, such as retraining a model with counterfactual data, representation projection, and prompting often fail to efficiently eliminate bias or directly alter the models' biased internal representations. To address these issues, we propose BiasEdit, an efficient model editing method to remove stereotypical bias from language models through lightweight networks that act as editors to generate parameter updates. BiasEdit employs a debiasing loss guiding editor networks to conduct local edits on partial parameters of a language model for debiasing while preserving the language modeling abilities during editing through a retention loss. Experiments on StereoSet and Crows-Pairs demonstrate the effectiveness, efficiency, and robustness of BiasEdit in eliminating bias compared to tangental debiasing baselines and little to no impact on the language models' general capabilities. In addition, we conduct bias tracing to probe bias in various modules and explore bias editing impacts on different components of language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。