arXiv:2508.02360cs.CL2025-08ACL被引 6

发现大模型政治立场会跨话题扩散,并提出抑制性微调方法缓解此问题。

Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models

  • 通过激活对比定位政治神经元,区分通用与特定话题神经元。
  • 抑制性微调使跨话题立场偏移降低20%,仅需抑制5%神经元。
  • 适用于关注模型公平性与可控性的研究者与应用开发者。

在政治话题上微调大型语言模型会显著改变其在多个议题上的政治立场,并意外影响无关话题的立场。现有研究虽指出该问题,但缺乏对内部表征及跨话题泛化机制的理解。本文从神经元层面系统探究该现象的内在机制并提出缓解方法。首先提出基于激活对比的政治神经元定位方法(PNLAC),识别出两类政治神经元:跨领域通用型神经元和单话题特异性神经元。在四个模型与数据集上通过激活修补实验验证了这两类神经元的存在。基于此,提出抑制性微调方法InhibitFT,有效缓解跨话题立场扩散。实验表明,所识别神经元类型在多种模型与数据集上具有鲁棒性,且InhibitFT平均将跨话题立场偏移降低20%,同时保持话题特定性能。此外,仅抑制5%的神经元即可有效缓解该问题。

原文摘要 · Abstract (English)

Fine-tuning Large Language Models on a political topic will significantly manipulate their political stance on various issues and unintentionally affect their stance on unrelated topics. While previous studies have proposed this issue, there is still a lack of understanding regarding the internal representations of these stances and the mechanisms that lead to unintended cross-topic generalization. In this paper, we systematically explore the internal mechanisms underlying this phenomenon from a neuron-level perspective and how to mitigate the cross-topic generalization of political fine-tuning. Firstly, we propose Political Neuron Localization through Activation Contrasting (PNLAC) to identify two distinct types of political neurons: general political neurons, which govern stance across multiple political topics, and topic-specific neurons} that affect the model's political stance on individual topics. We find the existence of these political neuron types across four models and datasets through activation patching experiments. Leveraging these insights, we introduce InhibitFT, an inhibition-based fine-tuning method, effectively mitigating the cross-topic stance generalization. Experimental results demonstrate the robustness of identified neuron types across various models and datasets, and show that InhibitFT significantly reduces the cross-topic stance generalization by 20% on average, while preserving topic-specific performance. Moreover, we demonstrate that selectively inhibiting only 5% of neurons is sufficient to effectively mitigate the cross-topic stance generalization.

大模型政治立场神经元分析微调优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。