arXiv:2605.30717cs.CL2026-05被引 2

通过神经元干预实现语言模型性别生成的精准控制。

Neuron-Level Interventions for Gendered and Gender-Neutral Generation in Language Models

论文配图:Neuron-Level Interventions for Gendered and Gender-Neutral Generation in Language Models
图 1 · 摘自论文原文
  • 定位并操控与性别相关的特定神经元,实现对生成内容的性别导向控制。
  • 实验显示早期层神经元集中编码性别信息,且干预后保持语义一致性。
  • 适用于性别偏见缓解与可控生成研究,适合关注公平性与可控性的开发者。

语言模型在中性提示下仍可能生成带有性别刻板印象的内容。现有研究多聚焦于男女二元对立,忽视了they/them等中性表达。本文研究了语言模型中与女性、男性及中性三类性别相关的神经元,提出一种基于神经元层级的干预方法,通过激活或屏蔽特定神经元,可有效引导生成目标性别形式,同时保持原意不变。我们构建了两个涵盖三类性别的控制句数据集,并通过人工评估验证质量。在两个开源模型上的实验表明,性别相关神经元主要集中在早期层,后期贡献较小。相比现有方法,本方法实现更精准的性别控制,非目标性别泄漏更低,输出质量更稳定。本工作揭示了性别在模型中的编码机制,为神经元干预评估与偏见缓解提供了有效工具。代码与数据集已公开。

原文摘要 · Abstract (English)

Language models (LMs) can produce gendered language and stereotypes even when given neutral prompts. Most prior work on gender bias in LMs primarily examines gender through a binary lens (feminine vs. masculine), with limited attention to gender-neutral forms, such as they/them pronouns or neutrally phrased job titles. How gender-related signals are encoded in the internal representations of LMs remains an open question. In this work, we study gender-specific neurons in LMs across three categories: feminine, masculine, and gender-neutral. We propose a neuron-level intervention method to identify neurons that are strongly tied to each gender category. We then test these neurons through controlled generation, showing that activating or masking gender-related neurons can steer a sentence toward a target gender form while preserving its original meaning. To evaluate the effectiveness of our gender-intervention approach, we curate two datasets with controlled sentences labeled across all three gender categories and validate the data quality through human evaluation. Experiments on two open-source LMs show that gender-specific neurons are not evenly distributed across model layers; instead, they concentrate heavily in the earliest layers with smaller contributions from later layers. Compared to existing methods, our method achieves more precise gender control, with less leakage into non-target gender categories and stable output quality through two evaluation criteria. Overall, our work examines how gender is encoded in LMs and provides a simple yet effective approach toward controlled gender intervention for both neuron intervention evaluation and gender bias mitigation. Code and datasets are available at: https://github.com/zhiwenyou103/Gender-Neuron-Intervention

语言模型性别偏见神经元干预可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。