构建72万张对抗样本数据集,测试图像生成检测器的鲁棒性。
RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors
- 用7个顶尖检测器对抗4种文生图模型生成对抗样本
- 攻击样本对未见检测器转移成功率高,验证鲁棒性缺陷
- 适合研究生成内容检测与安全防御的学者使用
AI生成图像已达到人类难以辨别真伪的程度。为应对欺诈与虚假信息风险,检测生成图像成为紧迫且活跃的研究课题。尽管诸多方法宣称具备高准确率,但多在理想条件下评估,常忽视对抗鲁棒性,或因分析复杂而被忽略。本文提出RAID(Robust evaluation of AI-generated image Detectors)数据集,包含72,000张多样且高度可迁移的对抗样本。该数据集通过攻击七个最先进的检测器,针对四种不同文生图模型生成的图像构造而成。大量实验表明,这些对抗样本能以高成功率转移到未见过的检测器,可快速提供对检测器鲁棒性的可靠估计。结果揭示当前顶尖生成图像检测器极易被对抗样本欺骗,凸显亟需更鲁棒的方法。数据集与评估代码已公开于https://huggingface.co/datasets/aimagelab/RAID和https://github.com/pralab/RAID。
原文摘要 · Abstract (English)
AI-generated images have reached a quality level at which humans are incapable of reliably distinguishing them from real images. To counteract the inherent risk of fraud and disinformation, the detection of AI-generated images is a pressing challenge and an active research topic. While many of the presented methods claim to achieve high detection accuracy, they are usually evaluated under idealized conditions. In particular, the adversarial robustness is often neglected, potentially due to a lack of awareness or the substantial effort required to conduct a comprehensive robustness analysis. In this work, we tackle this problem by providing a simpler means to assess the robustness of AI-generated image detectors. We present RAID (Robust evaluation of AI-generated image Detectors), a dataset of 72k diverse and highly transferable adversarial examples. The dataset is created by running attacks against an ensemble of seven state-of-the-art detectors and images generated by four different text-to-image models. Extensive experiments show that our methodology generates adversarial images that transfer with a high success rate to unseen detectors, which can be used to quickly provide an approximate yet still reliable estimate of a detector's adversarial robustness. Our findings indicate that current state-of-the-art AI-generated image detectors can be easily deceived by adversarial examples, highlighting the critical need for the development of more robust methods. We release our dataset at https://huggingface.co/datasets/aimagelab/RAID and evaluation code at https://github.com/pralab/RAID.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。