arXiv:2606.26285cs.CRcs.AI2026-06

提出可精准控制的扩散模型中毒攻击,隐蔽性强且效果显著。

TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models

论文配图:TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models
图 1 · 摘自论文原文
  • 将恶意扰动限制在时间维度上,保持输入分布一致
  • 支持多区域、多图像的特征重建,攻击成功率高
  • 适合研究生成数据安全与防御机制的研究者

基于噪声的扩散模型后门攻击通常依赖输入时触发注入、非目标激活和分布外目标生成,这些假设降低了攻击的隐蔽性和实际意义。本文提出TEMPO-Diffusion,一种定向后门框架,将恶意分布偏移局限于时间维度内的分布内暴露。该框架支持:(i) 针对特定类别的定向攻击;(ii) 多子图后门,可在多个不同输出图像中特定位置重建特定特征;(iii) 基于时间条件触发的修复生成。为研究利用带毒扩散模型生成合成训练数据带来的实际安全风险,我们还引入CALISA:一个平衡的、区域感知的交通标志数据集,强调加拿大和美国道路标志。在CIFAR10、GTSRB和CALISA上,实验表明TEMPO-Diffusion能可靠地污染特定类别合成数据生成,并使下游分类器在该数据上训练时达到高攻击成功率。

原文摘要 · Abstract (English)

Noise-based backdoor attacks on diffusion models typically rely on input-time trigger injection, untargeted activation, and out-of-distribution target generation. Such assumptions reduce both the stealthiness and the practical relevance of these attacks. In this work, we present TEMPO-Diffusion, a targeted backdoor framework that localizes the malicious distribution shift to a temporal, in-distribution exposure. TEMPO-Diffusion supports: (i) targeted attacks on and to specific classes, (ii) multiple sub-image backdoors that reconstruct specific features within multiple, different output images and at multiple locations, and (iii) in-painting with time-conditioned triggers. To study relevant, practical security concerns in leveraging backdoored diffusion models for synthetic training data, we also introduce CALISA: a balanced, region-aware traffic-sign dataset emphasizing Canadian and U.S. road signs. Across CIFAR10, GTSRB, and CALISA, our experiments show that TEMPO-Diffusion can reliably poison class-specific synthetic data generation and induce high attack success rates in downstream classifiers trained on that data.

后门攻击扩散模型生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。