arXiv:2509.25673cs.CL2025-09EMNLP被引 11

通过双路径遗忘机制,在不损失模型能力的前提下消除语言模型中的刻板印象。

Mitigating Biases in Language Models via Bias Unlearning

  • 设计双路径机制,分别遗忘刻板印象并保留反刻板内容。
  • 在多个评测基准上显著降低偏见,同时保持文本连贯性和任务准确率。
  • 方法可迁移至不同模型变体,揭示偏见在预训练中已固化。

大量研究发现语言模型对不同人口群体存在各类偏见,加剧歧视并损害公平性。现有参数修改类去偏方法严重损害文本连贯性和任务准确性;基于提示的去偏方法仅对预定义触发词有效,无法处理模型参数中深层嵌入的刻板关联。本文提出BiasUnlearn,一种新型模型去偏框架,通过双路径遗忘机制协同实现刻板印象遗忘与反刻板印象保留,并利用对抗遗忘集和动态数据集替换防止偏见极性反转。我们在多种语言模型和多个评估基准上进行了广泛实验,结果表明BiasUnlearn在缓解语言模型偏见的同时有效保留语言建模能力。进一步实验显示,去偏权重可在不同模型变体间迁移,证实偏见表示在预训练阶段即固化并贯穿微调过程。

原文摘要 · Abstract (English)

Many studies have shown various biases targeting different demographic groups in language models, amplifying discrimination and harming fairness. Recent parameter modification debiasing approaches significantly degrade core capabilities such as text coherence and task accuracy. And Prompt-based debiasing methods, only effective for predefined trigger words, fail to address deeply embedded stereotypical associations in model parameters. In this paper, we propose BiasUnlearn, a novel model debiasing framework which achieves targeted debiasing via dual-pathway unlearning mechanisms coordinating stereotype forgetting with anti-stereotype retention, while preventing bias polarity reversal through adversarial forget set and dynamic dataset swapping. We conducted extensive experiments with multiple language models across various evaluation benchmarks. The results show that BiasUnlearn outperforms existing methods in mitigating bias in language models while retaining language modeling capabilities. Further experiments reveal that debiasing weights are transferable across model variants, confirming that bias representations become entrenched during pre-training and persist through fine-tuning phases.

去偏语言模型刻板印象双路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。