扩散模型生成数据可能携带隐蔽后门,影响下游模型安全。
Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
- 利用扩散模型生成合成数据时,后门触发器会被记忆并传递。
- 不同攻击方式下,后门在合成数据中保持高有效性,且不影响数据质量。
- 细调策略配合特殊损失设计可低成本实现后门传播,适合研究者警惕。
扩散模型广泛用于下游感知任务的合成数据增强,显著降低数据收集与标注成本。然而,该新兴数据供应链存在严重安全隐患。公开可用的生成模型常未经验证即被复用,其安全性存疑。本文研究此类生成数据链中的后门传播问题——数据链后门(DCB)。发现开源扩散模型可成为后门的隐性载体,其强大的分布拟合能力导致模型记忆并再现后门触发器,进而被下游模型继承,造成严重安全风险。该威胁在无标签攻击场景下尤为突出,因对合成数据的实用性影响极小而难以察觉。我们对比了从头训练和微调两种攻击路径:直接微调效果弱,但通过优化损失函数与触发器处理机制,可显著提升触发器保留能力,使微调成为低成本攻击路径。在标准数据增强及数据稀缺设置下评估,多种触发类型均能稳定保留在合成数据中,攻击效果接近传统后门攻击。
原文摘要 · Abstract (English)
The increasing use of generative models such as diffusion models for synthetic data augmentation has greatly reduced the cost of data collection and labeling in downstream perception tasks. However, this new data source paradigm may introduce important security concerns. Publicly available generative models are often reused without verification, raising a fundamental question of their safety and trustworthiness. This work investigates backdoor propagation in such emerging generative data supply chain, namely, Data-Chain Backdoor (DCB). Specifically, we find that open-source diffusion models can become hidden carriers of backdoors. Their strong distribution-fitting ability causes them to memorize and reproduce backdoor triggers in generation, which are subsequently inherited by downstream models, resulting in severe security risks. This threat is particularly concerning under clean-label attack scenarios, as it remains effective while having negligible impact on the utility of the synthetic data. We study two attacker choices to obtain a backdoor-carried generator, training from scratch and fine-tuning. While naive fine-tuning leads to weak inheritance of the backdoor, we find that novel designs in the loss objectives and trigger processing can substantially improve the generator's ability to preserve trigger patterns, making fine-tuning a low-cost attack path. We evaluate the effectiveness of DCB under the standard augmentation protocol and further assess data-scarce settings. Across multiple trigger types, we observe that the trigger pattern can be consistently retained in the synthetic data with attack efficacy comparable to the conventional backdoor attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。