提升生成人脸检测系统抗对抗攻击能力
Robustness in AI-Generated Detection: Enhancing Resistance to Adversarial Attacks
- 结合对抗训练与扩散重建,增强检测鲁棒性
- 微小对抗扰动即可绕过现有系统,新方法显著提升防御能力
- 适合关注生成内容安全与模型防御的研究者
生成图像技术的快速发展带来了显著的安全隐患,尤其在人脸生成检测领域。本文研究了当前AI生成人脸检测系统的脆弱性。实验表明,尽管现有检测方法在标准条件下通常具有高准确率,但在对抗攻击下表现出有限的鲁棒性。为应对这一挑战,我们提出一种集成对抗训练、扩散反演与重构的方法,以缓解对抗样本的影响。实验结果表明,微小的对抗扰动即可轻易绕过现有检测系统,而我们的方法显著提升了系统的鲁棒性。此外,我们对对抗样本与正常样本进行了深入分析,揭示了AI生成内容的内在特征。所有相关代码将公开于专用仓库,以促进后续研究与验证。
原文摘要 · Abstract (English)
The rapid advancement of generative image technology has introduced significant security concerns, particularly in the domain of face generation detection. This paper investigates the vulnerabilities of current AI-generated face detection systems. Our study reveals that while existing detection methods often achieve high accuracy under standard conditions, they exhibit limited robustness against adversarial attacks. To address these challenges, we propose an approach that integrates adversarial training to mitigate the impact of adversarial examples. Furthermore, we utilize diffusion inversion and reconstruction to further enhance detection robustness. Experimental results demonstrate that minor adversarial perturbations can easily bypass existing detection systems, but our method significantly improves the robustness of these systems. Additionally, we provide an in-depth analysis of adversarial and benign examples, offering insights into the intrinsic characteristics of AI-generated content. All associated code will be made publicly available in a dedicated repository to facilitate further research and verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。