首次全面评测开源图像检测模型零样本性能,揭示其实际效果差异巨大。
How well are open sourced AI-generated image detection models out-of-the-box: A comprehensive benchmark study
- 测试16种主流检测器在23个预训练版本上的零样本表现
- 最佳检测器准确率75.0%,最差仅37.5%,差距达37个百分点
- 现代商业生成模型(如Midjourney v7)让多数检测器失效,平均准确率仅18-30%
随着AI生成图像在数字平台广泛传播,可靠的检测方法对防范虚假信息、保障内容真实性至关重要。尽管已有众多深度伪造检测方法提出,现有基准测试主要评估微调后的模型,忽视了实际部署中最常见的零样本场景。本文首次系统性地对16种前沿检测方法(共23个预训练变体)进行零样本评估,覆盖12个多样化数据集,包含260万至291种不同生成器的图像样本,涵盖现代扩散模型。分析发现:(1)无通用最优检测器,不同数据集间排名极不稳定(斯皮尔曼相关系数ρ:0.01–0.87);(2)最佳检测器均值准确率达75.0%,最差为37.5%,差距达37个百分点;(3)训练数据对齐显著影响泛化能力,同架构检测器间性能波动达20%–60%;(4)现代商业生成器(Flux Dev、Firefly v4、Midjourney v7)使多数检测器失效,平均准确率仅18%–30%;(5)识别出三类影响跨数据集泛化的系统性失败模式。统计检验显示检测器性能差异显著(Friedman检验:χ²=121.01,p<10⁻¹⁶,Kendall W=0.524)。研究挑战了“一劳永逸”检测器的理念,提出应根据具体威胁环境选择检测器,而非依赖公开基准结果。
原文摘要 · Abstract (English)
As AI-generated images proliferate across digital platforms, reliable detection methods have become critical for combating misinformation and maintaining content authenticity. While numerous deepfake detection methods have been proposed, existing benchmarks predominantly evaluate fine-tuned models, leaving a critical gap in understanding out-of-the-box performance -- the most common deployment scenario for practitioners. We present the first comprehensive zero-shot evaluation of 16 state-of-the-art detection methods, comprising 23 pretrained detector variants (due to multiple released versions of certain detectors), across 12 diverse datasets, comprising 2.6~million image samples spanning 291 unique generators including modern diffusion models. Our systematic analysis reveals striking findings: (1)~no universal winner exists, with detector rankings exhibiting substantial instability (Spearman~$ρ$: 0.01 -- 0.87 across dataset pairs); (2)~a 37~percentage-point performance gap separates the best detector (75.0\% mean accuracy) from the worst (37.5\%); (3)~training data alignment critically impacts generalization, causing up to 20--60\% performance variance within architecturally identical detector families; (4)~modern commercial generators (Flux~Dev, Firefly~v4, Midjourney~v7) defeat most detectors, achieving only 18--30\% average accuracy; and (5)~we identify three systematic failure patterns affecting cross-dataset generalization. Statistical analysis confirms significant performance differences between detectors (Friedman test: $χ^2$=121.01, $p<10^{-16}$, Kendall~$W$=0.524). Our findings challenge the ``one-size-fits-all'' detector paradigm and provide actionable deployment guidelines, demonstrating that practitioners must carefully select detectors based on their specific threat landscape rather than relying on published benchmark performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。