首个针对自监督扩散模型表征层的隐蔽后门攻击,可精准控制生成结果。
BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models
- 在PCA空间中劫持中毒样本语义,通过多空间约束操控去噪轨迹
- 在多个数据集上FID和MSE指标显著优于现有攻击方法
- 攻击隐蔽性强,正常功能不受影响,适合研究模型安全漏洞
自监督扩散模型通过潜在空间去噪学习高质量视觉表征。然而,其表征层存在独特威胁:不同于传统攻击生成输出,其无约束的潜在语义空间允许隐蔽后门,触发时可实现恶意控制。本文提出BadRSSD,首个针对自监督扩散模型表征层的后门攻击。具体而言,它将中毒样本在主成分分析(PCA)空间中的语义表示,劫持为目标图像的表示,再通过在潜在、像素和特征分布空间施加协同约束,控制扩散过程中的去噪轨迹,引导模型生成指定目标。此外,将表征分散正则化集成到约束框架中,保持特征空间均匀性,显著提升攻击隐蔽性。该方法在保持高模型实用性的同时,触发时实现高精度目标生成。在多个基准数据集上的实验表明,BadRSSD在FID和MSE指标上均显著优于现有攻击,在不同架构与配置下可靠建立后门,并有效抵御现有先进后门防御机制。
原文摘要 · Abstract (English)
Self-supervised diffusion models learn high-quality visual representations via latent space denoising. However, their representation layer poses a distinct threat: unlike traditional attacks targeting generative outputs, its unconstrained latent semantic space allows for stealthy backdoors, permitting malicious control upon triggering. In this paper, we propose BadRSSD, the first backdoor attack targeting the representation layer of self-supervised diffusion models. Specifically, it hijacks the semantic representations of poisoned samples with triggers in Principal Component Analysis (PCA) space toward those of a target image, then controls the denoising trajectory during diffusion by applying coordinated constraints across latent, pixel, and feature distribution spaces to steer the model toward generating the specified target. Additionally, we integrate representation dispersion regularization into the constraint framework to maintain feature space uniformity, significantly enhancing attack stealth. This approach preserves normal model functionality (high utility) while achieving precise target generation upon trigger activation (high specificity). Experiments on multiple benchmark datasets demonstrate that BadRSSD substantially outperforms existing attacks in both FID and MSE metrics, reliably establishing backdoors across different architectures and configurations, and effectively resisting state-of-the-art backdoor defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。