arXiv:2602.10382cs.CL2026-02被引 1

发现大模型后门触发器会劫持原有语言能力,而非新建电路。

Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models

  • 通过激活修补定位触发器,发现其依赖现有语言注意力头。
  • 触发器头与自然语言输出头重叠度达0.18~0.43(顶10头)。
  • 适合关注模型安全、可解释性防御的研究者阅读。

后门攻击对大型语言模型构成重大安全威胁,但触发器的内部工作机制仍不清晰。本文首次对预训练阶段注入的语言切换型后门进行机制分析,研究Gaperon模型族(1B、8B和24B)。利用激活修补技术,我们定位了触发器形成位置,并识别出处理触发器与自然语言信息的注意力头。核心发现是:触发器头与模型各规模下天然编码输出语言的注意力头存在显著重叠,顶10个头的雅各布森指数在0.18至0.43之间。这表明后门触发器并非构建新电路,而是劫持模型已有的语言组件与表征。该发现为后门检测与缓解策略提供了新思路,可借助触发器与自然行为的纠缠关系设计防御方法。本工作推动了对预训练注入型后门更真实、可解释的理解,为基于可解释性的防御奠定了基础。

原文摘要 · Abstract (English)

Backdoor attacks pose significant security risks for Large Language Models (LLMs), yet the internal mechanisms by which triggers operate remain poorly understood. We present the first mechanistic analysis of trigger-induced language-switching backdoors injected during pre-training, studying the Gaperon model family (1B, 8B and 24B). Using activation patching, we localize trigger formation and identify which attention heads process trigger and natural language information. Our central finding is that trigger heads substantially overlap with heads naturally encoding output language across model scales, with Jaccard indices between 0.18 and 0.43 over the top 10 heads identified. This suggests that backdoor triggers do not form new circuits but instead co-opt the model's existing language components and representations. These findings have implications for backdoor defense as detection methods and mitigation strategies could leverage this entanglement between triggers and natural behaviors. More broadly, our work represents a first step toward a more realistic mechanistic understanding of pre-training-injected backdoors in LLMs, paving the way for principled, interpretability-driven defenses.

后门攻击可解释性语言模型机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。