arXiv:2603.24125cs.CL2026-03被引 1

提出统一框架,发现对齐能减少输出偏见但不消除内部编码偏见。

Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study

  • 用相同中性提示同时分析模型内部表征与输出偏见
  • 对齐后输出偏见下降,但内部仍保留性别关联信息
  • 真实场景下去偏效果不稳健,对抗提示可唤醒隐藏偏见

大型语言模型在训练中会习得社会规律,导致下游应用中的性别偏见。现有缓解方法多聚焦于减少生成输出的偏见,通常在结构化基准上评估,存在两大问题:输出层面评估无法揭示对齐是否改变模型底层表征;结构化基准可能无法反映真实使用场景。本文提出统一框架,采用相同中性提示联合分析大模型的内在与外在性别偏见,实现直接比较内部表征中编码的性别信息与输出表达的偏见。与以往研究报道的弱或不一致相关性不同,我们在统一协议下发现潜在性别信息与表达偏见之间存在稳定关联。进一步通过监督微调进行对齐,结果显示虽能有效降低表达偏见,但内部表征中仍存在可测量的性别相关性,且在对抗性提示下可被重新激活。最后,我们考察了两种真实场景,发现结构化基准上观察到的去偏效果并不一定具备泛化能力,例如在故事生成任务中表现不佳。

原文摘要 · Abstract (English)

During training, Large Language Models (LLMs) learn social regularities that can lead to gender bias in downstream applications. Most mitigation efforts focus on reducing bias in generated outputs, typically evaluated on structured benchmarks, which raises two concerns: output-level evaluation does not reveal whether alignment modifies the model's underlying representations, and structured benchmarks may not reflect realistic usage scenarios. We propose a unified framework to jointly analyze intrinsic and extrinsic gender bias in LLMs using identical neutral prompts, enabling direct comparison between gender-related information encoded in internal representations and bias expressed in generated outputs. Contrary to prior work reporting weak or inconsistent correlations, we find a consistent association between latent gender information and expressed bias when measured under the unified protocol. We further examine the effect of alignment through supervised fine-tuning aimed at reducing gender bias. Our results suggest that while the latter indeed reduces expressed bias, measurable gender-related associations are still present in internal representations, and can be reactivated under adversarial prompting. Finally, we consider two realistic settings and show that debiasing effects observed on structured benchmarks do not necessarily generalize, e.g., to the case of story generation.

大模型偏见对齐机制表征分析去偏策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。