arXiv:2508.01525cs.CVcs.AI2025-08中稿 · ACMMM 2025被引 5

MiraGe提升识别未知AI生成图像能力,让检测更通用。

MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection

  • 通过特征对齐与类间分离,学习跨生成器的通用特征
  • 在多个基准上达到顶尖性能,对新生成器如Sora仍有效
  • 结合文本提示增强模型泛化能力,适合对抗新型生成模型

生成模型的快速发展凸显了区分真实图像与AI生成图像的鲁棒检测器的重要性。现有方法在已知生成器上表现良好,但在面对新出现或未见过的生成模型时性能显著下降,主要因特征嵌入重叠导致跨生成器分类不准确。本文提出多模态判别表征学习方法MiraGe,旨在学习生成器无关的特征表示。基于类内变化最小化与类间分离的理论启发,MiraGe紧密对齐同一类内的特征,同时最大化不同类间的分离度,从而提升特征判别性。此外,通过将多模态提示学习引入CLIP,利用文本嵌入作为语义锚点,进一步优化判别表征学习过程,增强泛化能力。在多个基准上的全面实验表明,MiraGe实现当前最优性能,在面对如Sora等未见生成器时依然保持强鲁棒性。

原文摘要 · Abstract (English)

Recent advances in generative models have highlighted the need for robust detectors capable of distinguishing real images from AI-generated images. While existing methods perform well on known generators, their performance often declines when tested with newly emerging or unseen generative models due to overlapping feature embeddings that hinder accurate cross-generator classification. In this paper, we propose Multimodal Discriminative Representation Learning for Generalizable AI-generated Image Detection (MiraGe), a method designed to learn generator-invariant features. Motivated by theoretical insights on intra-class variation minimization and inter-class separation, MiraGe tightly aligns features within the same class while maximizing separation between classes, enhancing feature discriminability. Moreover, we apply multimodal prompt learning to further refine these principles into CLIP, leveraging text embeddings as semantic anchors for effective discriminative representation learning, thereby improving generalizability. Comprehensive experiments across multiple benchmarks show that MiraGe achieves state-of-the-art performance, maintaining robustness even against unseen generators like Sora.

AI图像检测生成模型泛化能力多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。