arXiv:2511.18635cs.CLcs.AI2025-11被引 6

针对特定偏见的去偏技术可能加剧其他偏见,需多维度评估。

No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases

  • 在多个模型上测试四种去偏方法,覆盖种族、宗教、职业和性别偏见。
  • 针对性去偏常降低目标偏见,但会增加其他偏见并降低文本连贯性。
  • 提醒研究者:去偏需跨维度评估,避免偏见转移或恶化。

大型语言模型从训练数据中继承社会偏见,可能导致有害或不公平输出。尽管已有多种去偏技术,其效果通常仅在目标维度上评估。本文研究针对性去偏对跨类别偏见的影响。我们在七个模型族的十种模型上应用四种去偏技术,考察种族、宗教、职业与性别相关偏见。使用StereoSet基准衡量去偏对模型连贯性和刻板印象偏好影响。结果一致显示:虽然针对性去偏有时能减少目标维度偏见,但常导致其他维度偏见加剧,甚至降低整体连贯性。这凸显了在开发和评估去偏策略时,亟需具备多维度鲁棒性评估工具,以防无意中转移或恶化未被关注的偏见轴。

原文摘要 · Abstract (English)

Large Language Models (LLMs) inherit societal biases from their training data, potentially leading to harmful or unfair outputs. While various techniques aim to mitigate these biases, their effects are often evaluated only along the dimension of the bias being targeted. This work investigates the cross-category consequences of targeted bias mitigation. We study four bias mitigation techniques applied across ten models from seven model families, and we explore racial, religious, profession- and gender-related biases. We measure the impact of debiasing on model coherence and stereotypical preference using the StereoSet benchmark. Our results consistently show that while targeted mitigation can sometimes reduce bias in the intended dimension, it frequently leads to unintended and often negative consequences in others, such as increasing model bias and decreasing general coherence. These findings underscore the critical need for robust, multi-dimensional evaluation tools when examining and developing bias mitigation strategies to avoid inadvertently shifting or worsening bias along untargeted axes.

语言模型去偏偏见转移评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。