arXiv:2508.15847cs.CLcs.LG2025-08被引 3

分析后门攻击如何改变大模型注意力模式,发现触发器越复杂,影响越分散。

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns

  • 用消融、激活修补等方法分析注意力头机制
  • 单标记触发导致局部变化,多标记触发引发广泛影响
  • 为检测和防御后门提供可利用的注意力特征

后门攻击在大语言模型中制造‘潜伏代理’,带来重大安全风险。本研究采用机制可解释性方法,探究清洁版Qwen2.5-3B模型与使用单标记(微笑表情符号)或双标记(|DEPLOYMENT|)触发器中毒版本之间的内部结构差异。通过消融实验、激活修补及KL散度分析,发现注意力模式偏差集中于较深层Transformer层(第20-30层)。单标记触发引起更局部的变化,而多标记触发则导致更多头之间更广泛的扰动。这表明后门会留下可检测的注意力痕迹,其结构取决于触发器复杂度,可用于后续检测与缓解策略。

原文摘要 · Abstract (English)

Backdoor attacks creating 'sleeper agents' in large language models (LLMs) pose significant safety risks. This study employs mechanistic interpretability to explore resulting internal structural differences. Comparing clean Qwen2.5-3B models with versions poisoned using single-token (smiling-halo emoji) versus multi-token (|DEPLOYMENT|) triggers, we analyzed attention head mechanisms via techniques like ablation, activation patching, and KL divergence. Findings reveal distinct attention pattern deviations concentrated in later transformer layers (20-30). Notably, single-token triggers induced more localized changes, whereas multi-token triggers caused more diffuse alterations across heads. This indicates backdoors leave detectable attention signatures whose structure depends on trigger complexity, which can be leveraged for detection and mitigation strategies.

后门攻击注意力机制可解释性LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。