通过可解释神经元编辑,精准消除大模型性别偏见。
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
- 构建新数据集并定位导致偏见的特定神经元回路。
- 仅修改少量通用神经元即可显著降低偏见,且不损害模型能力。
- 适合关注AI公平性与模型可解释性的研究者与开发者。
大型语言模型常表现出性别偏见,影响其安全应用。现有方法或缺乏对偏见机制的深入理解,或损害模型核心能力。为此,我们提出 CommonWords 数据集,系统评估 LLM 中的性别偏见。分析发现偏见普遍存在,并识别出性别神经元与通用神经元等关键神经回路。值得注意的是,即使编辑少量通用神经元,也可能因层级交互破坏模型整体性能。基于此,我们提出一种结合逻辑值与因果策略的可解释神经元编辑方法,精准靶向偏见神经元。在五款 LLM 上的实验表明,该方法有效降低性别偏见,同时保持原有能力,优于现有微调与编辑方法。研究贡献包括新数据集、偏见机制分析及实用化解法。
原文摘要 · Abstract (English)
Large language models (LLMs) often exhibit gender bias, posing challenges for their safe deployment. Existing methods to mitigate bias lack a comprehensive understanding of its mechanisms or compromise the model's core capabilities. To address these issues, we propose the CommonWords dataset, to systematically evaluate gender bias in LLMs. Our analysis reveals pervasive bias across models and identifies specific neuron circuits, including gender neurons and general neurons, responsible for this behavior. Notably, editing even a small number of general neurons can disrupt the model's overall capabilities due to hierarchical neuron interactions. Based on these insights, we propose an interpretable neuron editing method that combines logit-based and causal-based strategies to selectively target biased neurons. Experiments on five LLMs demonstrate that our method effectively reduces gender bias while preserving the model's original capabilities, outperforming existing fine-tuning and editing approaches. Our findings contribute a novel dataset, a detailed analysis of bias mechanisms, and a practical solution for mitigating gender bias in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。