融合多维度特征提升生成图像检测准确率
Multi-Feature Fusion Approach for Generative AI Images Detection
- 结合低、中、高三个层次的视觉特征进行检测
- 在四个数据集上性能优于现有方法,混合模型下表现更稳
- 适合需要可靠检测能力的媒体审核与内容安全场景
生成式AI模型的快速发展催生了高度逼真的合成图像,传统方法难以有效区分其与真实照片。现有检测器多依赖单一特征空间(如统计规律、语义嵌入或纹理模式),在面对多样且不断演进的生成模型时鲁棒性不足。本文系统评估了一种多特征融合框架,整合三个互补特征空间:(1) 均值减去对比度归一化(MSCN)捕捉低层统计偏差;(2) CLIP嵌入编码高层语义一致性;(3) 多尺度局部二值模式(MLBP)表征中层纹理异常。在涵盖多种生成模型的四个基准数据集上,实验表明各单特征表现差异显著。关键在于三者融合后整体性能显著提升,尤其在跨模型混合场景下表现更优。相比先进方法,本框架在所有数据集上均实现一致改进。结果凸显混合表示对鲁棒生成图像检测的重要性,并提供一种整合互补视觉线索的可扩展框架。
原文摘要 · Abstract (English)
The rapid evolution of Generative AI (GenAI) models has led to synthetic images of unprecedented realism, challenging traditional methods for distinguishing them from natural photographs. While existing detectors often rely on single-feature spaces, such as statistical regularities, semantic embeddings, or texture patterns, these approaches tend to lack robustness when confronted with diverse and evolving generative models. In this work, we investigate and systematically evaluate a multi-feature fusion framework that combines complementary cues from three distinct spaces: (1) Mean Subtracted Contrast Normalized (MSCN) features capturing low-level statistical deviations; (2) CLIP embeddings encoding high-level semantic coherence; and (3) Multi-scale Local Binary Patterns (MLBP) characterizing mid-level texture anomalies. Through extensive experiments on four benchmark datasets covering a wide range of generative models, we show that individual feature spaces exhibit significant performance variability across different generators. Crucially, the fusion of all three representations yields superior and more consistent performance, particularly in a challenging mixed-model scenario. Compared to state-of-the-art methods, the proposed framework yields consistently improved performance across all evaluated datasets. Overall, this work highlights the importance of hybrid representations for robust GenAI image detection and provides a principled framework for integrating complementary visual cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。