通过消除模型内在偏见,提升大模型在决策中的公平性。
Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs
- 提出FACU方法,显式平衡刻板与反刻板关联的概率差异。
- 在多个数据集上显著降低性别偏见,且下游公平性普遍提升。
- 适合关注模型公平性、需兼顾性能与伦理的AI研发者使用。
大语言模型广泛应用于高风险决策系统,其偏见预测可能加剧社会经济不平等。尽管已有研究分别考察了内在表征偏见与下游不公平行为,但尚未明确缓解内在偏见是否能带来更公平的下游结果。本文提出公平感知概念去偏(FACU),一种面向公平性的模型级缓解方法,将概念去偏适配至表征平衡任务。不同于抑制型方法,FACU显式正则化刻板与反刻板关联之间的概率差异,同时保持预测性能和语言建模质量。我们在三个开源LLM上评估FACU,涵盖多个内在偏见基准和三个社会经济分类数据集,采用冻结的LLM嵌入与LoRA微调分类器两种方式。结果表明,FACU在多数设置下实现统计显著的内在性别偏见降低,并伴随下游公平性提升,且未明显损害预测性能。结合外部缓解方法(尤其是反事实数据增强)可进一步提升公平性。研究提示:公平导向的内在偏见缓解有助于实现更公平的LLM决策,偏见治理应贯穿模型开发与下游部署全过程。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities. Although prior work has examined intrinsic representational bias and unfair downstream behavior separately, it remains unclear whether mitigating intrinsic bias leads to fairer downstream outcomes. We introduce Fairness-Aware Concept Unlearning (FACU), a model-level mitigation method that adapts concept unlearning to fairness-oriented representation balancing. Unlike suppression-based approaches, FACU explicitly regularizes probability differences between stereotypical and anti-stereotypical associations while preserving predictive performance and language modeling quality. We evaluate FACU across three open-source LLMs, multiple intrinsic bias benchmarks, and three socio-economic classification datasets using both frozen LLM embeddings and LoRA-fine-tuned classifiers. FACU produces statistically significant reductions in intrinsic gender bias that are associated with downstream fairness improvements across most evaluated settings, datasets, models, and fairness metrics, without significantly degrading predictive performance. Combining FACU with extrinsic mitigation methods, particularly counterfactual data augmentation, yields further fairness improvements. These findings suggest that fairness-aware intrinsic mitigation can support fairer LLM-based decision-making and that bias mitigation should be addressed across both model development and downstream deployment stages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。