现有AI生成图像检测器易受真实世界攻击,鲁棒性堪忧。
Adversarial Robustness of AI-Generated Image Detectors in the Real World
- 用四种检测方法和五种攻击算法验证检测器在真实场景下的脆弱性。
- 攻击后分类性能显著下降,且在社交平台压缩后仍有效。
- 商业工具如HIVE也暴露同样问题,适合关注AI安全的研究者。
生成式人工智能(GenAI)的快速发展伴随着其滥用的严重风险,尤其是生成可信度高的虚假图像,对民主进程中的公众信任构成威胁。因此,迫切需要可靠区分真实与AI生成内容的工具。目前多数检测方法基于神经网络,通过识别伪造痕迹来判断。本文通过大量实验,涵盖四种检测方法和五种攻击算法,证明当前最先进的分类器在真实条件下极易受到对抗样本攻击,且攻击者无需了解检测器内部结构即可显著降低其性能。值得注意的是,大多数攻击在图像上传至社交媒体平台经历压缩等降质后依然有效。案例研究显示,商业工具如HIVE也存在类似鲁棒性缺陷。此外,我们评估了使用鲁棒预训练模型生成特征的方法,虽提升了一定抗性,但仍无法达到良性输入下的性能水平。这些结果结合GenAI侵蚀公众信任的潜在风险,凸显亟需新方法和新视角以防止其滥用。
原文摘要 · Abstract (English)
The rapid advancement of Generative Artificial Intelligence (GenAI) capabilities is accompanied by a concerning rise in its misuse. In particular the generation of credible misinformation in the form of images poses a significant threat to the public trust in democratic processes. Consequently, there is an urgent need to develop tools to reliably distinguish between authentic and AI-generated content. The majority of detection methods are based on neural networks that are trained to recognize forensic artifacts. In this work, we demonstrate that current state-of-the-art classifiers are vulnerable to adversarial examples under real-world conditions. Through extensive experiments, comprising four detection methods and five attack algorithms, we show that an attacker can dramatically decrease classification performance, without internal knowledge of the detector's architecture. Notably, most attacks remain effective even when images are degraded during the upload to, e.g., social media platforms. In a case study, we demonstrate that these robustness challenges are also found in commercial tools by conducting black-box attacks on HIVE, a proprietary online GenAI media detector. In addition, we evaluate the robustness of using generated features of a robust pre-trained model and showed that this increases the robustness, while not reaching the performance on benign inputs. These results, along with the increasing potential of GenAI to erode public trust, underscore the need for more research and new perspectives on methods to prevent its misuse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。