arXiv:2602.01283cs.CV2026-02被引 6

发现跨语言共用的安全神经元,可提升低资源语言的安全性。

Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons

  • 定位并验证了跨语言共享的安全神经元,是安全行为的关键调控者。
  • 抑制这些神经元导致低资源语言安全性集体下降,增强则提升跨语言防御一致性。
  • 仅微调少量神经元即可显著改善低资源语言安全表现,适合多语言安全优化场景。

多语言安全能力严重失衡,非高资源(NHR)语言远不如高资源(HR)语言安全。尽管已观察到跨语言表征转移现象,但其神经机制仍不明确。本文发现大模型中存在一组跨语言共享安全神经元(SS-Neurons),数量极少却对多语言安全行为起关键作用。首先识别出单语安全神经元(MS-Neurons),并通过定向激活与抑制验证其在拒绝安全请求中的因果作用。进一步分析发现,SS-Neurons是HR与NHR语言间共享的MS-Neurons子集,构成安全能力从高资源语言向低资源语言迁移的桥梁。实验显示,抑制这些神经元会导致多个NHR语言的安全性同步下降;而增强它们则能提升跨语言防御一致性。基于此,提出一种面向神经元的训练策略,根据语言资源分布和模型架构靶向优化SS-Neurons。实验证明,仅微调该极小神经元子集,即优于现有最先进方法,在显著提升NHR语言安全性的同时保持模型通用能力。代码与数据集将发布于https://github.com/1518630367/SS-Neuron-Expansion。

原文摘要 · Abstract (English)

Multilingual safety remains significantly imbalanced, leaving non-high-resource (NHR) languages vulnerable compared to robust high-resource (HR) ones. Moreover, the neural mechanisms driving safety alignment remain unclear despite observed cross-lingual representation transfer. In this paper, we find that LLMs contain a set of cross-lingual shared safety neurons (SS-Neurons), a remarkably small yet critical neuronal subset that jointly regulates safety behavior across languages. We first identify monolingual safety neurons (MS-Neurons) and validate their causal role in safety refusal behavior through targeted activation and suppression. Our cross-lingual analyses then identify SS-Neurons as the subset of MS-Neurons shared between HR and NHR languages, serving as a bridge to transfer safety capabilities from HR to NHR domains. We observe that suppressing these neurons causes concurrent safety drops across NHR languages, whereas reinforcing them improves cross-lingual defensive consistency. Building on these insights, we propose a simple neuron-oriented training strategy that targets SS-Neurons based on language resource distribution and model architecture. Experiments demonstrate that fine-tuning this tiny neuronal subset outperforms state-of-the-art methods, significantly enhancing NHR safety while maintaining the model's general capabilities. The code and dataset will be available athttps://github.com/1518630367/SS-Neuron-Expansion.

多语言安全对齐神经元分析低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。