arXiv:2607.28319cs.CLcs.CY2026-07

通过激活差异定位大模型中的人口统计偏见神经元。

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

论文配图:Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
图 1 · 摘自论文原文
  • 用对比提示捕捉推理时激活,识别GLU-MLP层中对人口属性敏感的神经元。
  • 仅零化40个神经元(不足总宽度0.031%)即保留99.49%模型能力。
  • 揭示偏见与能力可分离,为定向调控提供新方法,适合模型安全研究者。

本文提出公平性剪枝(Fairness Pruning),一种轻量级结构干预方法,用于管理和未来缓解大型语言模型(LLMs)中的种族/性别等人口统计偏见。基于因果偏见定位的实证验证,该方法采用极小差异的提示对与推理时激活捕获,识别在处理人口属性时反应不同的神经元,并评估其在down_proj输入处的信号。实验在参数量达30亿的模型(如Llama-3.2系列和Salamandra-2B)上进行,结合标准基准测试与定性文本生成。结果表明,零化这些神经元会改变模型对相关人口变量的响应。然而,因偏见评分无符号,候选集混合了推动和反向刻板印象的神经元,整体偏见净效应取决于主导符号。干预极为精准:在Llama-3.2-1B中最多零化40个神经元(<0.031%总MLP宽度),即可实现平均99.49%的推理与通用知识能力保留。这些发现实证确认了人口偏见处理与模型能力存在于可分离的神经回路中,为从盲目置零转向定向行为调节奠定了方法基础。

原文摘要 · Abstract (English)

This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs). As a foundational empirical validation of this method, this work focuses on causal bias localization. Using minimally contrastive prompt pairs and inference-time activation capture, the method identifies neurons that react differentially when processing demographic attributes in GLU architectures, evaluating the signal at the down_proj input. Empirical evaluation was conducted on models of up to 3 billion parameters (Llama-3.2 family and Salamandra-2B), combining standardized benchmark evaluation with qualitative text generation experiments. Results demonstrate that zeroing the identified neurons alters how the model responds to associated demographic variables. However, rather than producing flat mitigation, the intervention causes bidirectional bias destabilization: because BiasScore is unsigned, candidate sets mix neurons that push toward and against the stereotype, and the net effect on aggregate bias depends on which sign dominates. The intervention is extremely surgical: zeroing at most 40 neurons in Llama-3.2-1B (less than 0.031% of total MLP width) achieves a mean retention of 99.49% in reasoning and general knowledge capabilities. These findings empirically confirm that demographic bias processing and model capabilities operate on dissociable circuits, establishing the methodological foundations for transitioning from blind zeroing toward directional behavior modulation.

偏见检测神经元定位模型可解释性剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。