arXiv:2511.12046cs.CRcs.AI2025-11

用微小扰动即可实现知识蒸馏的隐蔽后门攻击

BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning

  • 无需代理模型,仅通过微调教师模型植入后门
  • 弱触发器在标准蒸馏中转移成功率高,攻击有效
  • 方法简单高效,比以往更隐蔽,适合安全研究者

知识蒸馏(KD)是压缩大模型的重要技术,但依赖第三方下载的预训练教师模型会带来严重安全风险,尤其是后门攻击。现有方法通常复杂且计算量大:使用代理学生模型和模拟蒸馏以保证迁移性,并构造类似通用对抗扰动的触发器,其幅度明显,具有强对抗性特征。本文质疑这种复杂性是否必要,提出一种隐蔽的“弱触发器”——感知上几乎不可察觉、对抗效应微弱的扰动。我们提出BackWeak,一种无代理模型的简单攻击范式。实验表明,仅用极小学习率对良性教师模型进行微调,加入弱触发器即可成功植入后门,该后门在受害者标准蒸馏过程中可可靠传递至多种学生架构,实现高攻击成功率。在多个数据集、模型结构和蒸馏方法上的大量实验证明,BackWeak效率更高、更简单,且常比以往复杂方法更具隐蔽性。本工作呼吁研究者关注触发器潜在的对抗特性。

原文摘要 · Abstract (English)

Knowledge Distillation (KD) is essential for compressing large models, yet relying on pre-trained "teacher" models downloaded from third-party repositories introduces serious security risks--most notably backdoor attacks. Existing KD backdoor methods are typically complex and computationally intensive: they employ surrogate student models and simulated distillation to guarantee transferability, and construct triggers similar to universal adversarial perturbations (UAPs), which being not stealthy in magnitude, inherently exhibit strong adversarial behavior. This work questions whether such complexity is necessary and constructs stealthy "weak" triggers--imperceptible perturbations that have negligible adversarial effect. We propose BackWeak, a simple, surrogate-free attack paradigm. BackWeak shows that a powerful backdoor can be implanted by simply fine-tuning a benign teacher with a weak trigger using a very small learning rate. We demonstrate that this delicate fine-tuning is sufficient to embed a backdoor that reliably transfers to diverse student architectures during a victim's standard distillation process, yielding high attack success rates. Extensive empirical evaluations on multiple datasets, model architectures, and KD methods show that BackWeak is efficient, simpler, and often more stealthy than previous elaborate approaches. This work calls on researchers studying KD backdoor attacks to pay particular attention to the trigger's potential adversarial characteristics.

后门攻击知识蒸馏隐蔽性安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。