arXiv:2608.09095cs.AI2026-08

发现跨语言安全能力的共享路径,仅调整少量参数即可提升低资源语言安全性。

Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

论文配图:Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways
图 1 · 摘自论文原文
  • 通过分析安全信号传播路径,定位跨语言共享的安全机制。
  • 仅更新少量路径参数,即显著提升非高资源语言的安全性。
  • 适合关注多语言AI安全与高效对齐方法的研究者。

揭示大语言模型安全能力的内在机制对于构建可信人工智能至关重要。当前关于多语言安全的可解释性研究多局限于局部组件(如孤立神经元),但这种静态、碎片化的视角忽略了组件间的协同作用,未能阐明安全信号如何在模型内部动态传播并最终驱动安全决策。本文突破孤立神经元的局限,识别并靶向安全信号传播过程中形成的跨层功能路径,从而揭示跨语言安全差距的成因。具体而言,我们首先识别单语言安全路径,并验证其对拒绝有害请求的影响;后续跨语言分析发现,存在一个稀疏的跨语言共享安全路径子集,证实该交集是将高资源(HR)语言的安全能力传递至非高资源(NHR)语言的内部桥梁。基于这些机制发现,我们提出一种基于跨语言共享安全路径的靶向对齐方法。实验表明,仅更新少量路径参数,即可显著提升NHR语言的安全性,同时基本保持模型的通用能力。

原文摘要 · Abstract (English)

Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely confined to local components, such as isolated neurons. However, this static and fragmented perspective overlooks the synergy among components and fails to elucidate how safety signals dynamically propagate within the model to drive safety decisions ultimately. In this work, we move beyond isolated neurons to identify and target the cross-layer functional pathways formed during safety signal propagation, thereby uncovering the mechanisms driving the cross-lingual safety gap. Specifically, we first identify monolingual safety pathways and validate their impact on refusing harmful requests. Subsequent cross-lingual analyses reveal a sparse subset of cross-lingual shared safety pathways, confirming that this intersection acts as the internal bridge transferring safety capabilities from high-resource (HR) languages to non-high-resource (NHR) languages. Building on these mechanistic findings, we propose a pathways-targeted alignment method based on the cross-lingual shared safety pathways. Experimental results show that updating only a small fraction of pathway parameters significantly improves safety in NHR languages while largely preserving the model's general capabilities.

多语言安全可解释性路径靶向对齐方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。