arXiv:2602.18880cs.CVcs.AI2026-02

用多模态大模型融合空间与频域特征,实现伪造图像的精准检测与可解释定位。

FOCA: Frequency-Oriented Cross-Domain Forgery Detection, Localization and Explanation via Multi-Modal Large Language Model

  • 融合RGB空间与频率域特征,通过交叉注意力机制提升检测能力。
  • 在新构建的FSE-Set数据集上,检测准确率超越现有方法。
  • 提供跨域可解释性分析,适合需要透明推理的媒体验证场景。

生成式模型等图像篡改技术的发展对媒体真实性验证、数字取证和公众信任构成重大挑战。现有图像伪造检测与定位(IFDL)方法存在两大缺陷:过度依赖语义内容而忽视纹理线索;对细微低层篡改痕迹的可解释性不足。为此,我们提出FOCA——一种基于多模态大语言模型的框架,通过交叉注意力融合模块整合来自RGB空间域和频率域的判别特征,实现高精度伪造检测与定位,并提供明确的人类可读的跨域解释。我们还构建了FSE-Set,一个包含多样化真实与篡改图像、像素级掩码及双域标注的大规模数据集。大量实验表明,FOCA在空间与频率域上的检测性能与可解释性均优于当前最先进方法。

原文摘要 · Abstract (English)

Advances in image tampering techniques, particularly generative models, pose significant challenges to media verification, digital forensics, and public trust. Existing image forgery detection and localization (IFDL) methods suffer from two key limitations: over-reliance on semantic content while neglecting textural cues, and limited interpretability of subtle low-level tampering traces. To address these issues, we propose FOCA, a multimodal large language model-based framework that integrates discriminative features from both the RGB spatial and frequency domains via a cross-attention fusion module. This design enables accurate forgery detection and localization while providing explicit, human-interpretable cross-domain explanations. We further introduce FSE-Set, a large-scale dataset with diverse authentic and tampered images, pixel-level masks, and dual-domain annotations. Extensive experiments show that FOCA outperforms state-of-the-art methods in detection performance and interpretability across both spatial and frequency domains.

图像伪造多模态可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。