用冻结视觉编码器检测AI生成图像,仅需1万张训练图仍很有效
SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders

- 利用冻结多模态编码器的嵌入空间区分真实与合成图像
- 仅用1万张图像训练,对未见生成器的鲁棒性优于大规模数据集
- 适合关注轻量级、高泛化能力检测方案的研究者
生成模型的快速发展模糊了合成图像与真实图像的界限,亟需可靠的深度伪造检测方法。现有方法大多依赖大规模真实-虚假图像数据集,但随着新生成器不断涌现,这类数据集越来越难维护。本文研究现代多模态视觉表示中是否已蕴含图像真实性信息。发现冻结的多模态编码器在嵌入空间中自然区分真实与合成图像,使无需任务微调的线性分类器即可取得良好性能。基于此,我们提出一种表征感知的数据筛选策略,仅选取10,000张代表性生成图像用于训练,相比AIGIBench的288,000张和OpenFake的4,000,000张显著减少。该方法提升了对未见生成器和分布偏移的鲁棒性。此外,我们构建了RealWorldBench基准,包含现代相机照片、当代图库图像及近期商业生成器输出。跨多个基准的实验表明,结合冻结多模态表示与精心筛选的训练数据,可实现简单高效的AI生成图像检测。
原文摘要 · Abstract (English)
The rapid advancement of generative models has blurred the boundary between synthetic and real imagery, creating an urgent need for reliable deepfake detection. Yet most existing approaches rely on massive real--fake datasets, which are increasingly difficult to maintain as new generators continue to emerge. In this work, we investigate how much information about image authenticity is already encoded in modern multimodal vision representations. We find that frozen multimodal encoders naturally separate real and synthetic images in their embedding space, enabling a simple linear classifier to achieve strong performance without task-specific fine-tuning. Motivated by this observation, we develop a representation-aware data curation strategy that selects a compact set of representative generators for training. The resulting training set contains only 10K images, compared to 288K in AIGIBench and 4M in OpenFake, while improving robustness to unseen generators and distribution shifts. We additionally introduce RealWorldBench, a benchmark consisting of modern camera photographs, contemporary stock images, and outputs from recent commercial generators. Experiments across multiple benchmarks show that combining frozen multimodal representations with carefully curated training data provides a simple and effective approach to AI-generated image detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。