提出可迁移的后门攻击方法,让模型在知识蒸馏中悄悄植入隐蔽后门。
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
- 设计复合触发词,利用常见词汇组合实现隐蔽后门
- 在4个大模型家族中验证后门可成功转移到学生模型
- 适用于研究知识蒸馏安全性的研究人员或安全工程师
大型语言模型常被下游用户用作教师模型进行知识蒸馏,以压缩为更高效的模型。然而,这些教师模型可能来自不可信来源,蒸馏过程可能带来意外安全风险。本文研究了从被污染教师模型中进行知识蒸馏的安全问题。首先发现,现有后门大多无法传递到学生模型,原因是其触发词在正常语境中极少出现。我们提出新方法T-MTB,构建由多个在目标蒸馏数据集中频繁出现的词组成的复合触发器,使中毒教师保持隐蔽,而蒸馏过程中各词的独立存在足以传递后门信号。通过T-MTB,我们在越狱和内容操控两种攻击场景下,对四个大模型家族进行了广泛实验,充分揭示了可迁移后门的安全风险。
原文摘要 · Abstract (English)
LLMs are often used by downstream users as teacher models for knowledge distillation, compressing their capabilities into memory-efficient models. However, as these teacher models may stem from untrusted parties, distillation can raise unexpected security risks. In this paper, we investigate the security implications of knowledge distillation from backdoored teacher models. First, we show that prior backdoors mostly do not transfer onto student models. Our key insight is that this is because existing LLM backdooring methods choose trigger tokens that rarely occur in usual contexts. We argue that this underestimates the security risks of knowledge distillation and introduce a new backdooring technique, T-MTB, that enables the construction and study of transferable backdoors. T-MTB carefully constructs a composite backdoor trigger, made up of several specific tokens that often occur individually in anticipated distillation datasets. As such, the poisoned teacher remains stealthy, while during distillation the individual presence of these tokens provides enough signal for the backdoor to transfer onto the student. Using T-MTB, we demonstrate and extensively study the security risks of transferable backdoors across two attack scenarios, jailbreaking and content modulation, and across four model families of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。