arXiv:2412.06727cs.CV2024-12被引 4

用真实后处理生成骗过AI内容检测的对抗样本

Take Fake as Real: Realistic-like Robust Black-box Adversarial Attack to Evade AIGC Detection

  • 用高斯模糊等真实后处理构造对抗样本,模拟自然图像特征
  • 在多种检测器上实现超70%欺骗率,且视觉几乎无损
  • 适合研究检测系统漏洞或提升生成内容隐蔽性的研究人员

AI生成内容(AIGC)检测的安全性对保障多媒体可信度至关重要。现有对抗攻击多聚焦于GAN生成人脸,难以应对多类自然图像及基于扩散模型的检测器,且隐蔽性差。本文深入分析检测器对不同后处理的脆弱性差异,发现检测器对特定后处理更敏感。针对实际场景中检测器不可知的情况,提出一种基于后处理融合优化的鲁棒黑盒攻击R²BA。不同于传统扰动,R²BA采用真实世界后处理(如高斯模糊、JPEG压缩、高斯噪声、光斑)生成对抗样本。通过带惯性衰减的随机粒子群算法优化后处理融合强度,依据检测器输出的伪造概率动态调整脆弱/鲁棒后处理强度,平衡攻击效果与视觉不可见性。在主流商用检测器和数据集上的实验表明,R²BA在GAN与扩散模型场景下均表现出优异的逃逸性能、良好隐蔽性和强鲁棒性。相比最先进的白盒与黑盒攻击,在原始与鲁棒场景下分别提升15%-72%和21%-47%的欺骗率,为现实应用中AIGC检测安全提供了重要启示。

原文摘要 · Abstract (English)

The security of AI-generated content (AIGC) detection is crucial for ensuring multimedia content credibility. To enhance detector security, research on adversarial attacks has become essential. However, most existing adversarial attacks focus only on GAN-generated facial images detection, struggle to be effective on multi-class natural images and diffusion-based detectors, and exhibit poor invisibility. To fill this gap, we first conduct an in-depth analysis of the vulnerability of AIGC detectors and discover the feature that detectors vary in vulnerability to different post-processing. Then, considering that the detector is agnostic in real-world scenarios and given this discovery, we propose a Realistic-like Robust Black-box Adversarial attack (R$^2$BA) with post-processing fusion optimization. Unlike typical perturbations, R$^2$BA uses real-world post-processing, i.e., Gaussian blur, JPEG compression, Gaussian noise and light spot to generate adversarial examples. Specifically, we use a stochastic particle swarm algorithm with inertia decay to optimize post-processing fusion intensity and explore the detector's decision boundary. Guided by the detector's fake probability, R$^2$BA enhances/weakens the detector-vulnerable/detector-robust post-processing intensity to strike a balance between adversariality and invisibility. Extensive experiments on popular/commercial AIGC detectors and datasets demonstrate that R$^2$BA exhibits impressive anti-detection performance, excellent invisibility, and strong robustness in GAN-based and diffusion-based cases. Compared to state-of-the-art white-box and black-box attacks, R$^2$BA shows significant improvements of 15\%--72\% and 21\%--47\% in anti-detection performance under the original and robust scenario respectively, offering valuable insights for the security of AIGC detection in real-world applications.

对抗攻击AIGC检测黑盒攻击后处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。