arXiv:2605.14091cs.CV2026-05被引 5

统一检测与定位各类伪造图像,性能超越39项基准测试。

Venus-DeFakerOne: Unified Fake Image Detection & Localization

论文配图:Venus-DeFakerOne: Unified Fake Image Detection & Localization
图 1 · 摘自论文原文
  • 融合InternVL2与SAM2构建统一模型,支持跨场景检测与像素级定位。
  • 在39个检测与9个定位基准上达到顶尖水平,对真实干扰和先进生成器鲁棒。
  • 揭示数据规模、跨域干扰等设计规律,适合安全监控与内容审核场景。

近年来,生成式AI的快速发展彻底改变了图像伪造范式,打破了文档编辑、自然图像处理、DeepFake生成与全图AIGC合成之间的传统界限。然而,现有的假图像检测与定位(FIDL)研究仍呈碎片化状态,导致统一伪造生成与领域特定检测之间的不匹配。为应对这一挑战,我们提出DeFakerOne,一个以数据为中心的统一FIDL基础模型,整合InternVL2与SAM2。该模型可实现跨多样化场景的图像级检测与像素级伪造定位。大量实验表明,DeFakerOne在39个伪造检测基准和9个定位基准上均优于现有基线。此外,模型对真实世界扰动及前沿生成器(如GPT-Image-2)表现出优异鲁棒性。最后,我们系统分析了数据缩放规律、跨域伪影传递与干扰模式、细粒度监督必要性以及原始分辨率伪影保留问题,揭示了可扩展、鲁棒且统一的FIDL设计原则。

原文摘要 · Abstract (English)

In recent years, the rapid evolution of generative AI has fundamentally reshaped the paradigm of image forgery, breaking the traditional boundaries between document editing, natural image manipulation, DeepFake generation, and full-image AIGC synthesis. Despite this shift toward unified forgery generation, existing research in Fake Image Detection and Localization (FIDL) remains fragmented. This creates a mismatch between increasingly unified forgery generation mechanisms and the domain-specific detection paradigm. Bridging this mismatch poses two key challenges for FIDL: understanding cross-domain artifacts transfer and interference, and building a high-capacity unified foundation model for joint detection and localization. To address these challenges, we propose DeFakerOne, a data-centric, unified FIDL foundation model integrating InternVL2 and SAM2. DeFakerOne enables simultaneous image-level detection and pixel-level forgery localization across diverse scenarios. Extensive experiments demonstrate that DeFakerOne achieves state-of-the-art performance, outperforming baselines on 39 forgery detection benchmarks and 9 localization benchmarks. Furthermore, the model exhibits superior robustness against real-world perturbations and state-of-the-art generators such as GPT-Image-2. Finally, we provide a systematic analysis of data scaling laws, cross-domain artifacts transfer-interference patterns, the necessity of fine-grained supervision, and the original resolution artifacts preservation, highlighting the design principles for scalable, robust, and unified FIDL.

伪造检测统一模型图像定位AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。