重建类假图检测器易受对抗攻击,伪造图像可骗过检测
Training-Free Reconstruction-Based AI-Generated Image Detectors Are Inherently Vulnerable to Adversarial Examples
- 针对自编码器重建误差设计对抗攻击,让假图伪装成真图
- 三种生成器、三类检测器均被攻破,性能大幅下降
- 攻击样本可跨检测器通用,暴露该类方法根本缺陷
AI生成图像的视觉质量与广泛存在,催生了可靠且鲁棒的检测方法需求。重建类检测器因其透明性和无需训练的特性,成为有前景的合成图像识别方向。然而,由于其运行机制与传统分类器显著不同,其对抗鲁棒性尚不明确。本文提出两种针对基于自编码器重建误差的检测器的新攻击方法。通过构建难以察觉的对抗样本,可人为增大原始图像与重建结果间的距离,导致假图被错误分类为真实图像。在三种顶尖生成器和三种检测器上的评估表明,即使攻击后图像经历真实世界退化,检测性能仍显著下降。关键发现是,这些对抗样本可自然迁移至不同检测器,因它们共享相同原理,揭示了重建类检测器的内在脆弱性。
原文摘要 · Abstract (English)
The impressive visual quality and ubiquity of AI-generated images call for reliable and robust detection methods. Reconstruction-based detectors have emerged as a promising direction for transparent and training-free identification of synthetic images. However, due to their fundamentally different mode of operation (compared to standard, classifier-based methods), little is known about their adversarial robustness. In this work, we propose two novel attack methods targeted at detectors that leverage autoencoder reconstruction error. We find that by constructing imperceptible adversarial examples, the distance between original and reconstruction can be artificially increased, causing fake images to be wrongly classified as real. Our evaluation including images from three state-of-the-art generators and three detectors demonstrates that detection performance is significantly decreased, even if attacked images additionally undergo real-world degradations. Critically, our adversarial examples naturally transfer across detectors, as they all share the same principle, pointing towards an inherent vulnerability of reconstruction-based detectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。