构建真实场景数据集,评估图像生成检测模型在复杂环境下的表现。
Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios
- 设计涵盖七类场景的鲁棒性数据集,覆盖真实世界多样性。
- 发现现有检测方法在跨平台传播和重数字化后性能显著下降。
- 揭示人类在少量样本下具备较强适应能力,可启发更鲁棒的检测系统。
随着生成模型快速发展,高度逼真的图像合成对数字安全与媒体可信度构成新挑战。尽管已有部分检测方法应对这一问题,但在复杂真实场景下的评估仍存在显著空白。本文提出真实世界鲁棒性数据集(RRDataset),从三方面全面评估检测模型:1)场景泛化性:涵盖战争冲突、灾害事故、政治社会事件、医疗公共健康、文化宗教、劳动生产及日常生活共七类高质图像,填补内容维度的数据缺口;2)互联网传输鲁棒性:测试图像经多个社交媒体平台多轮传播后的检测效果;3)重数字化鲁棒性:评估模型在四种不同重数字化方法处理后的表现。我们在该数据集上对17种检测器和10个视觉语言模型进行基准测试,并开展大规模人工实验,涉及192名参与者,研究人类在少量样本下的少样本学习能力。结果揭示当前AI检测方法在真实条件下存在明显局限,强调借鉴人类适应性以开发更鲁棒检测算法的重要性。
原文摘要 · Abstract (English)
With the rapid advancement of generative models, highly realistic image synthesis has posed new challenges to digital security and media credibility. Although AI-generated image detection methods have partially addressed these concerns, a substantial research gap remains in evaluating their performance under complex real-world conditions. This paper introduces the Real-World Robustness Dataset (RRDataset) for comprehensive evaluation of detection models across three dimensions: 1) Scenario Generalization: RRDataset encompasses high-quality images from seven major scenarios (War and Conflict, Disasters and Accidents, Political and Social Events, Medical and Public Health, Culture and Religion, Labor and Production, and everyday life), addressing existing dataset gaps from a content perspective. 2) Internet Transmission Robustness: examining detector performance on images that have undergone multiple rounds of sharing across various social media platforms. 3) Re-digitization Robustness: assessing model effectiveness on images altered through four distinct re-digitization methods. We benchmarked 17 detectors and 10 vision-language models (VLMs) on RRDataset and conducted a large-scale human study involving 192 participants to investigate human few-shot learning capabilities in detecting AI-generated images. The benchmarking results reveal the limitations of current AI detection methods under real-world conditions and underscore the importance of drawing on human adaptability to develop more robust detection algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。