arXiv:2603.09772cs.CVcs.CR2026-03

发现隐藏的后门触发器,现有防御可能失效。

Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors

  • 通过对比正常与触发样本特征,定位后门方向。
  • 新触发器可绕过防御,激活同一后门。
  • 应防御特征空间中的后门方向而非输入触发器。

当前后门防御假设消除已知触发器即可移除后门。我们证明这一视角不完整:存在与训练触发器在感知上差异明显的替代触发器,能可靠激活相同后门。通过对比干净样本与触发样本的特征表示,估计后门在特征空间的方向,并提出一种联合优化目标预测与方向对齐的特征引导攻击。理论上证明了替代触发器的存在是后门训练的必然结果,并实证验证了其有效性。许多仅移除训练触发器的防御仍留下后门,而替代触发器可利用潜藏的特征空间后门。研究呼吁防御应聚焦于表示空间的后门方向,而非输入空间的触发器。

原文摘要 · Abstract (English)

Current backdoor defenses assume that neutralizing a known trigger removes the backdoor. We show this trigger-centric view is incomplete: \emph{alternative triggers}, patterns perceptually distinct from training triggers, reliably activate the same backdoor. We estimate the alternative trigger backdoor direction in feature space by contrasting clean and triggered representations, and then develop a feature-guided attack that jointly optimizes target prediction and directional alignment. First, we theoretically prove that alternative triggers exist and are an inevitable consequence of backdoor training. Then, we verify this empirically. Additionally, defenses that remove training triggers often leave backdoors intact, and alternative triggers can exploit the latent backdoor feature-space. Our findings motivate defenses targeting backdoor directions in representation space rather than input-space triggers.

后门攻击特征空间防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。