arXiv:2508.01248cs.CV2025-08被引 10

通过分离CLIP语义信息,提升对未知生成图像的检测泛化能力。

NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection

  • 利用空域投影剥离CLIP特征中的语义信息,聚焦图像内在分布差异。
  • 在40种生成模型上实现7.4%的准确率提升,显著优于现有方法。
  • 适合需要跨模型检测生成图像的研究者与安全应用开发者。

生成模型(如GANs和扩散模型)的快速发展催生了高度逼真的图像,引发其在安全敏感领域被滥用的担忧。现有检测器在已知生成设置下表现良好,但在面对未知生成模型时泛化能力差,尤其当真实与伪造图像语义高度相似时。本文重新审视CLIP特征在生成图像检测中的应用,发现其高层语义信息阻碍有效区分。为此,提出NS-Net框架,通过空域投影将语义信息从CLIP视觉特征中解耦,并采用对比学习捕捉真实与生成图像间的内在分布差异。此外,设计局部块选择策略,通过抑制全局结构带来的语义偏差,保留细粒度伪影。在包含40种不同生成模型的开放世界基准上进行广泛实验,NS-Net优于现有最先进方法,检测准确率提升7.4%,展现出对GAN与扩散模型生成图像的强大泛化能力。

原文摘要 · Abstract (English)

The rapid progress of generative models, such as GANs and diffusion models, has facilitated the creation of highly realistic images, raising growing concerns over their misuse in security-sensitive domains. While existing detectors perform well under known generative settings, they often fail to generalize to unknown generative models, especially when semantic content between real and fake images is closely aligned. In this paper, we revisit the use of CLIP features for AI-generated image detection and uncover a critical limitation: the high-level semantic information embedded in CLIP's visual features hinders effective discrimination. To address this, we propose NS-Net, a novel detection framework that leverages NULL-Space projection to decouple semantic information from CLIP's visual features, followed by contrastive learning to capture intrinsic distributional differences between real and generated images. Furthermore, we design a Patch Selection strategy to preserve fine-grained artifacts by mitigating semantic bias caused by global image structures. Extensive experiments on an open-world benchmark comprising images generated by 40 diverse generative models show that NS-Net outperforms existing state-of-the-art methods, achieving a 7.4\% improvement in detection accuracy, thereby demonstrating strong generalization across both GAN- and diffusion-based image generation techniques.

图像检测CLIP生成模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。