通过优化触发器实现隐蔽性与攻击力兼得的干净图像后门攻击
Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization
- 用条件InfoGAN寻找自然图像中可作隐蔽触发器的特征
- 仅需极少量中毒样本即可攻击,准确率下降不足1%
- 适配多种数据集模型任务,对现有防御有较强抵抗力
干净图像后门攻击仅通过训练数据标签篡改即可危害深度神经网络,威胁关键应用安全。现有方法存在毒化率过高导致干净准确率(CA)显著下降的问题,削弱隐蔽性。本文提出生成式干净图像后门(GCB)框架,利用条件InfoGAN识别可作为强效且隐蔽触发器的自然图像特征。通过确保这些触发器易于与正常任务特征分离,使受害者模型仅从极少数中毒样本中即可学习后门,实现CA下降小于1%。实验表明GCB具有卓越泛化能力,成功应用于六个数据集、五个架构和四种任务,首次在回归与分割任务中实现干净图像后门攻击。同时,对多数现有后门防御表现出强鲁棒性。
原文摘要 · Abstract (English)
Clean-image backdoor attacks, which use only label manipulation in training datasets to compromise deep neural networks, pose a significant threat to security-critical applications. A critical flaw in existing methods is that the poison rate required for a successful attack induces a proportional, and thus noticeable, drop in Clean Accuracy (CA), undermining their stealthiness. This paper presents a new paradigm for clean-image attacks that minimizes this accuracy degradation by optimizing the trigger itself. We introduce Generative Clean-Image Backdoors (GCB), a framework that uses a conditional InfoGAN to identify naturally occurring image features that can serve as potent and stealthy triggers. By ensuring these triggers are easily separable from benign task-related features, GCB enables a victim model to learn the backdoor from an extremely small set of poisoned examples, resulting in a CA drop of less than 1%. Our experiments demonstrate GCB's remarkable versatility, successfully adapting to six datasets, five architectures, and four tasks, including the first demonstration of clean-image backdoors in regression and segmentation. GCB also exhibits resilience against most of the existing backdoor defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。