arXiv:2608.29936cs.CLcs.LG2026-08

揭示大模型安全与语言身份的内在纠缠机制,解释跨语言安全为何不统一。

When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs

论文配图:When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs
图 1 · 摘自论文原文
  • 用稀疏自编码器分析模型残差流中的安全特征分布
  • 发现安全特征与语言身份高度纠缠,影响跨语言安全表现
  • 为多语言安全干预提供可解释的机制依据,适合模型安全研究者

大型语言模型(LLMs)的安全对齐在不同语言间表现不一,但其内部机制尚不明确。本文通过稀疏自编码器(SAE)特征,系统分析了三个指令微调的LLM在八种语言、所有模型层中与有害和无害行为相关联的稀疏可解释方向。结果表明,安全相关特征的位置和分布具有架构依赖性;这些特征在几何上与语言身份纠缠,并表现出跨语言共享模式——不同语言在模型深度和架构上的安全特征共享程度各异。消融安全特征不仅影响有害响应率,还影响目标语言表现,干预效果可由安全特征与语言特征的关系预测。研究揭示安全对齐的语言普适性依赖于架构设计,为多语言安全干预提供了机制解释。

原文摘要 · Abstract (English)

Safety alignment of large language models (LLMs) degrades across languages, yet the internal mechanism driving this asymmetry remains poorly understood. Our work, therefore, presents a systematic mechanistic analysis of multilingual safety using sparse autoencoder (SAE) features, sparse interpretable directions in the residual stream associated with harmful and harmless model behavior across three instruction-tuned LLMs, eight languages, and all model layers. We observe that safety-relevant features are architecture-dependent in terms of where they are located and how they are distributed across layers. Additionally, they are geometrically entangled with language identity and exhibit cross-lingual sharing patterns, i.e., languages share safety features to varying degrees across model depths and architectures. This safety-language entanglement has direct consequences such that ablating safety features impacts not only harmful response rates but also target language, with the degree of intervention predicted by the relationship between safety and language features. Our findings qualify the language-universality of safety alignment as architecture-dependent and offer a mechanistic account of multilingual safety interventions.

安全对齐多语言机制分析语言纠缠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。