arXiv:2601.05663cs.SEcs.LG2026-01被引 1

发现并抑制预训练模型中的偏见神经元,提升软件工程应用的公平性

Tracing Stereotypes in Pre-trained Transformers: From Biased Neurons to Fairer Models

  • 通过构建九类偏见三元组数据集,定位模型中编码刻板印象的神经元
  • 抑制特定神经元后偏见显著降低,且在软件工程任务中性能损失极小
  • 为可解释的模型公平性改进提供新路径,适合关注AI伦理的研究者

基于Transformer的语言模型已广泛应用于软件工程领域,但其可能复制或放大社会偏见,引发公平性问题。现有研究显示,可通过神经元编辑修改模型内部激活以改变行为。本文基于知识神经元概念,假设预训练模型中存在编码刻板印象的偏见神经元。为此,我们构建了一个包含九类偏见关系的三元组数据集,并改编神经元归因策略,在BERT模型中追踪并抑制这些偏见神经元。实验表明,偏见知识集中在少数神经元中,抑制它们可显著降低偏见,同时在软件工程任务上仅造成微小性能损失。这证明了在神经元层面实现偏见溯源与缓解的可行性,为软件工程中的模型公平性提供了可解释的解决方案。

原文摘要 · Abstract (English)

The advent of transformer-based language models has reshaped how AI systems process and generate text. In software engineering (SE), these models now support diverse activities, accelerating automation and decision-making. Yet, evidence shows that these models can reproduce or amplify social biases, raising fairness concerns. Recent work on neuron editing has shown that internal activations in pre-trained transformers can be traced and modified to alter model behavior. Building on the concept of knowledge neurons, neurons that encode factual information, we hypothesize the existence of biased neurons that capture stereotypical associations within pre-trained transformers. To test this hypothesis, we build a dataset of biased relations, i.e., triplets encoding stereotypes across nine bias types, and adapt neuron attribution strategies to trace and suppress biased neurons in BERT models. We then assess the impact of suppression on SE tasks. Our findings show that biased knowledge is localized within small neuron subsets, and suppressing them substantially reduces bias with minimal performance loss. This demonstrates that bias in transformers can be traced and mitigated at the neuron level, offering an interpretable approach to fairness in SE.

模型公平性神经元编辑偏见检测BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。