简单线性分类器在视觉基础模型上实现高效AIGI检测。
Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models
- 用冻结的视觉基础模型特征+简单线性分类器,无需复杂结构。
- 在真实场景中检测准确率提升超30%,优于专门设计的检测器。
- 适合关注真实世界AIGI检测可靠性的研究者与应用开发者。
尽管针对AI生成图像(AIGI)的专用检测器在标准基准上接近完美,但在真实、非受控场景中性能急剧下降。本文证明:简单优于复杂。仅用现代视觉基础模型(如Perception Encoder、MetaCLIP 2、DINOv3)的冻结特征,训练一个简单线性分类器,即可达到新SOTA。在传统基准、未见过的生成器及挑战性真实分布上全面评估显示,该基线不仅在标准数据集上媲美专用检测器,更在真实场景中显著超越,准确率提升超过30%。我们认为这种能力是大规模预训练数据中包含合成内容所催生的涌现特性。其来源为两类数据暴露:视觉-语言模型内化了伪造的语义概念,自监督学习模型则从预训练数据中隐式获取了判别性取证特征。但模型仍存在局限:在图像重采样与传输后性能下降,对VAE重构和局部编辑仍无感知。我们主张将AI取证范式从静态基准过拟合,转向利用基础模型不断演化的世界知识以提升真实可靠性。
原文摘要 · Abstract (English)
While specialized detectors for AI-Generated Images (AIGI) achieve near-perfect accuracy on curated benchmarks, they suffer from a dramatic performance collapse in realistic, in-the-wild scenarios. In this work, we demonstrate that simplicity prevails over complex architectural designs. A simple linear classifier trained on the frozen features of modern Vision Foundation Models , including Perception Encoder, MetaCLIP 2, and DINOv3, establishes a new state-of-the-art. Through a comprehensive evaluation spanning traditional benchmarks, unseen generators, and challenging in-the-wild distributions, we show that this baseline not only matches specialized detectors on standard benchmarks but also decisively outperforms them on in-the-wild datasets, boosting accuracy by striking margins of over 30\%. We posit that this superior capability is an emergent property driven by the massive scale of pre-training data containing synthetic content. We trace the source of this capability to two distinct manifestations of data exposure: Vision-Language Models internalize an explicit semantic concept of forgery, while Self-Supervised Learning models implicitly acquire discriminative forensic features from the pretraining data. However, we also reveal persistent limitations: these models suffer from performance degradation under recapture and transmission, remain blind to VAE reconstruction and localized editing. We conclude by advocating for a paradigm shift in AI forensics, moving from overfitting on static benchmarks to harnessing the evolving world knowledge of foundation models for real-world reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。