arXiv:2607.07937cs.CL2026-07ACL

去偏处理可能引发意外偏见,反而加剧其他群体的刻板印象。

When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation

论文配图:When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation
图 1 · 摘自论文原文
  • 通过删减或替换群体指代词来去偏,但会引入新偏见。
  • 在维基百科数据上测试,多种方法均导致非目标群体偏见上升。
  • 注意力分布变化小,难以解释机制,需更透明的评估方式。

基于预处理的去偏方法(如在去偏语料库上预训练或微调)广泛用于自然语言处理。尽管这些方法能降低对特定群体的可测量偏见,但我们发现它们常引发意外的偏见转移——即对其他无关群体的刻板印象或反刻板印象反而增强,相对于中性基准。我们在两种模型架构(仅编码器和仅解码器)、多种预处理策略(删除刻板语句、移除群体提及、替换群体指代)以及不同数据规模下,在维基百科上验证了这一现象。标准评测基准往往无法捕捉这些偏见转移。通过注意力可视化分析,我们发现此类副作用并不伴随显著的注意力流变化,增加了机制解释的难度。论文讨论了评估体系的局限性,提出可操作的诊断工具,并呼吁采用考虑副作用的透明化去偏实践。

原文摘要 · Abstract (English)

Preprocessing-based methods for stereotype mitigation, such as pre-/post-training on debiased corpora, are widely used in NLP. While these approaches reduce measurable stereotypes for targeted groups, we find they often induce unintended shifts-side effects, where stereotyping or counter-stereotyping can increase relative to neutral baselines for other demographics, including across unrelated demographic categories. We demonstrate these side effects across two model families (encoder-only and decoder-only), multiple preprocessing strategies (removing stereotypical sentences, removing group mentions, and swapping group references), and both pre- and post-training at different data scales on Wikipedia. Standard benchmarks frequently miss these shifts. Using attention-rollout analysis, we observe that such side effects are not accompanied by large changes in attention flow, complicating mechanistic explanations. We discuss implications for evaluation, provide actionable diagnostics, and argue for side-effect-aware, transparent mitigation practices.

去偏刻板印象副作用NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。