用物理特征检测假图,跨模型效果好,还能提升大模型可信度。
Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection
- 提取5个稳定物理特征,跨数据集和生成模型有效区分真伪图像。
- 在多个基准上达99.8%准确率,显著优于现有方法。
- 结合CLIP模型,增强视觉语言理解,减少大模型幻觉问题。
人工智能生成内容(AIGC)快速发展,模糊了真实与合成图像的界限,暴露了现有深度伪造检测器对特定生成模型过拟合的问题。本文聚焦两个核心问题:(1) 哪些物理特征能在不同数据集和生成架构下稳定、鲁棒地区分自然图像与AI生成图像?(2) 这些像素级客观特征能否融入如CLIP等多模态模型中,提升检测性能并缓解基于语言信息的不可靠性?为此,我们在20余个由GAN与扩散模型生成的数据集上,系统评估了15种物理特征。提出一种新型特征选择算法,识别出拉普拉斯方差、Sobel统计量、残差噪声方差等五个核心特征,在所有测试数据集中均表现出一致的判别能力。这些特征被转化为文本编码值,并与语义描述一起用于引导CLIP中的图像-文本表征学习。大量实验表明,该方法在多个Genimage基准上达到顶尖性能,如在Wukong和SDv1.4数据集上准确率达99.8%。本工作首次将物理基础特征融入可信视觉语言建模,为缓解大模型幻觉与文本不准确问题开辟新路径。
原文摘要 · Abstract (English)
The rapid advancement of AI generated content (AIGC) has blurred the boundaries between real and synthetic images, exposing the limitations of existing deepfake detectors that often overfit to specific generative models. This adaptability crisis calls for a fundamental reexamination of the intrinsic physical characteristics that distinguish natural from AI-generated images. In this paper, we address two critical research questions: (1) What physical features can stably and robustly discriminate AI generated images across diverse datasets and generative architectures? (2) Can these objective pixel-level features be integrated into multimodal models like CLIP to enhance detection performance while mitigating the unreliability of language-based information? To answer these questions, we conduct a comprehensive exploration of 15 physical features across more than 20 datasets generated by various GANs and diffusion models. We propose a novel feature selection algorithm that identifies five core physical features including Laplacian variance, Sobel statistics, and residual noise variance that exhibit consistent discriminative power across all tested datasets. These features are then converted into text encoded values and integrated with semantic captions to guide image text representation learning in CLIP. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple Genimage benchmarks, with near-perfect accuracy (99.8%) on datasets such as Wukong and SDv1.4. By bridging pixel level authenticity with semantic understanding, this work pioneers the use of physically grounded features for trustworthy vision language modeling and opens new directions for mitigating hallucinations and textual inaccuracies in large multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。