流模型易受后门攻击,可高效实现精准图像篡改。
TrojFlow: Flow Models are Natural Targets for Trojan Attacks
- 将后门攻击建模为分布迁移,利用流模型强拟合能力实现隐蔽攻击
- 在CIFAR-10和CelebA上实现高保真、高特异性攻击,成功率超90%
- 揭示现有扩散模型防御手段对流模型攻击无效,适合安全评估研究者
基于流的生成模型(FMs)因其高效的训练与采样过程,在多个领域广泛应用,本质上可视为扩散模型(DMs)的一种变体。已有研究表明,扩散模型易受后门攻击,即通过输入中嵌入恶意模式触发输出篡改。本文发现,生成模型的后门攻击本质是将后门分布的图像迁移到目标分布的任务,而流模型具备拟合任意两分布的独特能力,极大简化了攻击训练与采样流程,使其天然成为后门攻击的理想目标。为此,本文提出TrojFlow,系统探索流模型的脆弱性。我们考察多种攻击设置及其组合,并评估现有扩散模型防御方法在这些场景下的有效性。在CIFAR-10与CelebA数据集上的实验表明,该方法可在保持高生成质量的同时实现高针对性攻击,且能轻易突破现有防御机制。
原文摘要 · Abstract (English)
Flow-based generative models (FMs) have rapidly advanced as a method for mapping noise to data, its efficient training and sampling process makes it widely applicable in various fields. FMs can be viewed as a variant of diffusion models (DMs). At the same time, previous studies have shown that DMs are vulnerable to Trojan/Backdoor attacks, a type of output manipulation attack triggered by a maliciously embedded pattern at model input. We found that Trojan attacks on generative models are essentially equivalent to image transfer tasks from the backdoor distribution to the target distribution, the unique ability of FMs to fit any two arbitrary distributions significantly simplifies the training and sampling setups for attacking FMs, making them inherently natural targets for backdoor attacks. In this paper, we propose TrojFlow, exploring the vulnerabilities of FMs through Trojan attacks. In particular, we consider various attack settings and their combinations and thoroughly explore whether existing defense methods for DMs can effectively defend against our proposed attack scenarios. We evaluate TrojFlow on CIFAR-10 and CelebA datasets, our experiments show that our method can compromise FMs with high utility and specificity, and can easily break through existing defense mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。