构建16万张图像基准测试,评估AI生成图像检测能力并提出视觉真实度指数。
The Visual Counter Turing Test (VCT2): A Benchmark for Evaluating AI-Generated Image Detection and the Visual AI Index (VAI)
- 设计VCT2基准,覆盖6大主流文生图模型输出的图像对。
- 零样本测试下检测准确率仅58%~58.34%,暴露现有方法泛化不足。
- 提出VAI指标,通过12个低层特征量化生成图像真实感,适合研究真实度与检测关系。
文本到图像生成模型的快速发展引发了对虚假视觉信息滥用的担忧。现有检测方法常过度拟合已知生成器,在新模型输出上表现不佳。本文提出视觉对抗图灵测试(VCT2),包含16.6万张图像的基准数据集,涵盖六种先进文生图系统:Stable Diffusion 2.1、SDXL、SD3 Medium、SD3.5 Large、DALL.E 3 和 Midjourney 6。数据集分为两类:基于MS COCO结构化描述的COCOAI,以及来自《纽约时报》推文的叙事风格TwitterAI。在统一的零样本评估下,17个主流检测模型在COCOAI和TwitterAI上的准确率分别为58%和58.34%,令人警觉。为突破二分类局限,提出视觉人工智能指数(VAI),一种基于十二个低层视觉特征的可解释、提示无关真实感度量,可更精细地量化与排序生成图像的感知质量。相关性分析显示VAI与检测准确率呈中等负相关(COCOAI: Pearson -0.532;TwitterAI: -0.503),表明越逼真的图像越难被检测,该趋势在各生成器间一致。代码与数据集均已公开,以推动通用化检测与真实感评估的发展。
原文摘要 · Abstract (English)
The rapid progress and widespread availability of text-to-image (T2I) generative models have heightened concerns about the misuse of AI-generated visuals, particularly in the context of misinformation campaigns. Existing AI-generated image detection (AGID) methods often overfit to known generators and falter on outputs from newer or unseen models. We introduce the Visual Counter Turing Test (VCT2), a comprehensive benchmark of 166,000 images, comprising both real and synthetic prompt-image pairs produced by six state-of-the-art T2I systems: Stable Diffusion 2.1, SDXL, SD3 Medium, SD3.5 Large, DALL.E 3, and Midjourney 6. We curate two distinct subsets: COCOAI, featuring structured captions from MS COCO, and TwitterAI, containing narrative-style tweets from The New York Times. Under a unified zero-shot evaluation, we benchmark 17 leading AGID models and observe alarmingly low detection accuracy, 58% on COCOAI and 58.34% on TwitterAI. To transcend binary classification, we propose the Visual AI Index (VAI), an interpretable, prompt-agnostic realism metric based on twelve low-level visual features, enabling us to quantify and rank the perceptual quality of generated outputs with greater nuance. Correlation analysis reveals a moderate inverse relationship between VAI and detection accuracy: Pearson of -0.532 on COCOAI and -0.503 on TwitterAI, suggesting that more visually realistic images tend to be harder to detect, a trend observed consistently across generators. We release COCOAI, TwitterAI, and all codes to catalyze future advances in generalized AGID and perceptual realism assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。