研究提示词长短对生成图像检测难易的影响,发现细节越丰富的提示生成的图越容易被识破。
Human vs. AI: A Novel Benchmark and a Comparative Study on the Detection of Generated Images and the Impact of Prompts
- 构建新数据集COCOXGEN,对比长短提示生成的图像检测效果。
- 长提示生成的图像被人类和AI检测器识别率均显著提升。
- 人类与AI关注的图像特征不同,揭示检测机制差异。
随着公开可用的文本到图像AI系统的发展,生成逼真但完全合成的图像已变得高度普及,可能通过简化虚假信息传播威胁公众。机器检测器和人类媒体素养有助于区分AI生成(虚假)图像与真实图像,以应对这一风险。尽管生成模型高度依赖提示词,但提示词对检测性能的影响尚未被充分研究。本工作考察提示词详细程度对伪造图像可检测性的影响,分别通过AI检测器和200名参与者的用户研究进行评估。为此,我们构建了新数据集COCOXGEN,包含来自COCO数据集的真实照片,以及使用SDXL和Fooocus在两种标准化长度提示下生成的图像。用户研究显示,使用更长、更详细的提示生成的图像被更易识别;同样,基于AI的检测模型在长提示生成图像上表现更好。然而,热力图分析表明,人类与AI模型关注的细节不同。
原文摘要 · Abstract (English)
With the advent of publicly available AI-based text-to-image systems, the process of creating photorealistic but fully synthetic images has been largely democratized. This can pose a threat to the public through a simplified spread of disinformation. Machine detectors and human media expertise can help to differentiate between AI-generated (fake) and real images and counteract this danger. Although AI generation models are highly prompt-dependent, the impact of the prompt on the fake detection performance has rarely been investigated yet. This work therefore examines the influence of the prompt's level of detail on the detectability of fake images, both with an AI detector and in a user study. For this purpose, we create a novel dataset, COCOXGEN, which consists of real photos from the COCO dataset as well as images generated with SDXL and Fooocus using prompts of two standardized lengths. Our user study with 200 participants shows that images generated with longer, more detailed prompts are detected significantly more easily than those generated with short prompts. Similarly, an AI-based detection model achieves better performance on images generated with longer prompts. However, humans and AI models seem to pay attention to different details, as we show in a heat map analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。