arXiv:2502.07771cs.CLcs.AI2025-02被引 5

通过剪枝分析大模型种族偏见,发现通用缓解策略效果有限。

Breaking Down Bias: On The Limits of Generalizable Pruning Strategies

  • 用神经元剪枝降低偏见,比剪枝注意力头更有效
  • 针对金融决策的剪枝在商业场景中效果迅速下降
  • 偏见具有强情境性,通用方案难奏效,适合关注法律责任者

我们采用模型剪枝来探究大语言模型如何表征种族偏见,并检验通用缓解策略是否可行。研究发现,剪枝可在不显著增加异常行为的前提下有效降低偏见;基于神经元的剪枝策略优于剪枝整个注意力头的方法。然而,随着剪枝策略泛化程度提高,其有效性迅速下降:在金融决策中训练的去偏模型无法有效应对商业交易中的偏见。整体表明,种族偏见在语言模型中仅部分以通用概念存在,其余高度依赖具体语境,因此通用缓解策略效果有限。该发现对AI法律框架有重要启示,提示应将特定应用场景中的模型部署责任明确分配给相关方。

原文摘要 · Abstract (English)

We employ model pruning to examine how LLMs conceptualize racial biases, and whether a generalizable mitigation strategy for such biases appears feasible. Our analysis yields several novel insights. We find that pruning can be an effective method to reduce bias without significantly increasing anomalous model behavior. Neuron-based pruning strategies generally yield better results than approaches pruning entire attention heads. However, our results also show that the effectiveness of either approach quickly deteriorates as pruning strategies become more generalized. For instance, a model that is trained on removing racial biases in the context of financial decision-making poorly generalizes to biases in commercial transactions. Overall, our analysis suggests that racial biases are only partially represented as a general concept within language models. The other part of these biases is highly context-specific, suggesting that generalizable mitigation strategies may be of limited effectiveness. Our findings have important implications for legal frameworks surrounding AI. In particular, they suggest that an effective mitigation strategy should include the allocation of legal responsibility on those that deploy models in a specific use case.

偏见缓解模型剪枝大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。