arXiv:2510.05740cs.CVcs.AI2025-10

提出双轴泛化框架与FusionDetect,提升假图检测在跨生成器和跨视觉域的表现。

Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect

  • 融合CLIP与Dinov2双模型特征,构建适应性强的统一特征空间
  • 在主流基准上准确率领先对手3.87%,在OmniGen上提升4.48%
  • 适用于跨领域、抗图像扰动的通用假图检测场景

生成模型的快速发展使可靠检测合成图像变得愈发重要。尽管现有研究多聚焦于跨生成器泛化,我们指出这视角过于局限。假图检测还面临另一关键挑战:跨视觉域泛化。为此,我们提出OmniGen基准数据集,涵盖12个先进生成器,更真实评估检测器性能。同时引入FusionDetect方法,利用两个冻结的基础模型(CLIP与Dinov2)的互补特征,构建能自然适应生成内容与设计变化的统一特征空间。大量实验表明,FusionDetect不仅达到新最优水平——在主流基准上比最接近的竞品高3.87%准确率,平均精度高6.13%,且在OmniGen上实现4.48%的准确率提升,并对常见图像扰动表现出卓越鲁棒性。本工作不仅提供高性能检测器,还推出新基准与框架,推动通用AI图像检测发展。代码与数据集已开源。

原文摘要 · Abstract (English)

The rapid development of generative models has made it increasingly crucial to develop detectors that can reliably detect synthetic images. Although most of the work has now focused on cross-generator generalization, we argue that this viewpoint is too limited. Detecting synthetic images involves another equally important challenge: generalization across visual domains. To bridge this gap,we present the OmniGen Benchmark. This comprehensive evaluation dataset incorporates 12 state-of-the-art generators, providing a more realistic way of evaluating detector performance under realistic conditions. In addition, we introduce a new method, FusionDetect, aimed at addressing both vectors of generalization. FusionDetect draws on the benefits of two frozen foundation models: CLIP & Dinov2. By deriving features from both complementary models,we develop a cohesive feature space that naturally adapts to changes in both thecontent and design of the generator. Our extensive experiments demonstrate that FusionDetect delivers not only a new state-of-the-art, which is 3.87% more accurate than its closest competitor and 6.13% more precise on average on established benchmarks, but also achieves a 4.48% increase in accuracy on OmniGen,along with exceptional robustness to common image perturbations. We introduce not only a top-performing detector, but also a new benchmark and framework for furthering universal AI image detection. The code and dataset are available at http://github.com/amir-aman/FusionDetect

假图检测跨域泛化多模态融合基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。