重建类检测器对扩散生成图像易受不可察觉攻击,准确率几乎归零。
Fragile Reconstruction: Adversarial Vulnerability of Reconstruction-Based Detectors for Diffusion-Generated Images
- 通过白盒攻击使三类主流检测器性能全面下降
- 攻击具迁移性,黑盒攻击亦可实现
- 现有防御方法效果有限,因样本信噪比过低
近期,针对扩散模型生成图像的检测受到广泛关注,因其可能带来的安全威胁。现有方法中,基于重建的检测范式尤为突出。然而,我们发现此类方法对对抗扰动极为脆弱:在输入图像中添加不可察觉的对抗扰动后,检测分类器的准确率几乎降至零。为验证此风险,我们在四种不同生成主干模型上系统评估了三种代表性检测器的对抗鲁棒性。首先,在白盒场景下构建对抗攻击,导致所有训练良好的检测器性能崩溃。此外,攻击表现出强迁移性——针对某一检测器设计的攻击可有效转移至其他检测器,表明黑盒攻击同样可行。最后,我们测试了常见防御措施,发现标准对抗防御方法仅能提供有限缓解。我们归因于攻击样本在检测器眼中信噪比过低。总体而言,结果揭示了基于重建的检测器存在根本性安全缺陷,亟需重新思考现有检测策略。
原文摘要 · Abstract (English)
Recently, detecting AI-generated images produced by diffusion-based models has attracted increasing attention due to their potential threat to safety. Among existing approaches, reconstruction-based methods have emerged as a prominent paradigm for this task. However, we find that such methods exhibit severe security vulnerabilities to adversarial perturbations; that is, by adding imperceptible adversarial perturbations to input images, the detection accuracy of classifiers collapses to near zero. To verify this threat, we present a systematic evaluation of the adversarial robustness of three representative detectors across four diverse generative backbone models. First, we construct adversarial attacks in white-box scenarios, which degrade the performance of all well-trained detectors. Moreover, we find that these attacks demonstrate transferability; specifically, attacks crafted against one detector can be transferred to others, indicating that adversarial attacks on detectors can also be constructed in a black-box setting. Finally, we assess common countermeasures and find that standard defense methods against adversarial attacks provide limited mitigation. We attribute these failures to the low signal-to-noise ratio (SNR) of attacked samples as perceived by the detectors. Overall, our results reveal fundamental security limitations of reconstruction-based detectors and highlight the need to rethink existing detection strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。