arXiv:2604.26772cs.CV2026-04中稿 · IEEE/CVF Conferenc…

用视觉大模型特征检测AI生成图像,性能比CLIP提升12%以上。

TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection

论文配图:TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection
图 1 · 摘自论文原文
  • 用可调注意力池化重设计分类头,融合现代视觉模型特征。
  • 在多个数据集上超越CLIP,最高提升12.3%准确率。
  • 适合做AI图像取证、生成内容检测的研究者和工程师。

近期方法表明,像CLIP视觉变换器这类大规模预训练模型作为特征提取器,能有效检测未见过的生成模型产生的AI生成图像(AIGIs)。当前最先进的AIGI检测方法多基于原始CLIP-ViT进行改进。自CLIP发布以来,涌现出众多视觉基础模型(VFMs),具备架构优化与不同训练范式。尽管如此,它们在AIGI检测与AI图像鉴证中的潜力仍鲜有探索。本文构建了跨多种VFM家族的全面基准,涵盖多样预训练目标、输入分辨率与模型规模,系统评估其对全生成及修补型AI图像的零样本检测性能。结果发现,最佳模型在准确率上超过原始CLIP逾12%,并优于已有方法。为充分挖掘现代VFM特征,我们提出简单重设计分类头的方法——可调注意力池化(TAP),将输出标记聚合为更优全局表征。结合最新VFMs使用TAP,在多个AIGI检测基准上实现显著提升,于两个具有挑战性的野外检测基准上建立新最优成绩。

原文摘要 · Abstract (English)

Recent methods demonstrate that large-scale pretrained models, such as CLIP vision transformers, effectively detect AI-generated images (AIGIs) from unseen generative models when used as feature extractors. Many state-of-the-art methods for AI-generated image detection build upon the original CLIP-ViT to enhance this generalization. Since CLIP's release, numerous vision foundation models (VFMs) have emerged, incorporating architectural improvements and different training paradigms. Despite these advances, their potential for AIGI detection and AI image forensics remains largely unexplored. In this work, we present a comprehensive benchmark across multiple VFM families, covering diverse pretraining objectives, input resolutions, and model scales. We systematically evaluate their out-of-the-box performance for detecting fully-generated AI-images and AI-inpainted images, and discover that the best model outperforms the original CLIP by more than 12% in accuracy, beating established approaches in the process. To fully leverage the features of a modern VFM, we propose a simple redesign of the classifier head by utilizing tunable attention pooling (TAP), which aggregates output tokens into a refined global representation. Integrating TAP with the latest VFMs yields substantial performance gains across several AIGI detection benchmarks, establishing a new state-of-the-art on two challenging benchmarks for in-the-wild detection of AI-generated and -inpainted images.

图像检测视觉模型AI取证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。