arXiv:2510.12229cs.CLcs.AI2025-10被引 4

发现微调后大模型会习得道德偏见,且可定位到特定层并修复。

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

  • 通过层替换分析,定位道德偏见出现的具体模型层。
  • 仅替换少数关键层的激活值,即可消除道德偏见。
  • 为大模型偏见干预提供无需重训练的新路径,适合安全研究者。

大语言模型在微调过程中可能内化人类类似的偏见,但其具体机制尚不明确。本文研究了著名的诺贝效应(Knobe effect)——一种在意图判断中的道德偏见——是否出现在微调后的大型语言模型中,并探究其在模型内部的定位方式。通过对3个开源权重的LLM进行层替换分析,我们发现该偏见不仅在微调中被学习,而且集中存在于特定的若干层中。令人意外的是,将预训练模型对应层的激活值替换到这些关键层,即可有效消除该效应。结果表明,大模型中的社会偏见可通过有针对性的干预实现解释、定位和缓解,而无需重新训练模型。

原文摘要 · Abstract (English)

Large language models (LLMs) have been shown to internalize human-like biases during finetuning, yet the mechanisms by which these biases manifest remain unclear. In this work, we investigated whether the well-known Knobe effect, a moral bias in intentionality judgements, emerges in finetuned LLMs and whether it can be traced back to specific components of the model. We conducted a Layer-Patching analysis across 3 open-weights LLMs and demonstrated that the bias is not only learned during finetuning but also localized in a specific set of layers. Surprisingly, we found that patching activations from the corresponding pretrained model into just a few critical layers is sufficient to eliminate the effect. Our findings offer new evidence that social biases in LLMs can be interpreted, localized, and mitigated through targeted interventions, without the need for model retraining.

模型偏见可解释性微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。