arXiv:2602.10079cs.CV2026-02

一个模型同时检测图像拼接与复制粘贴伪造,精准定位篡改区域和来源。

Can Image Splicing and Copy-Move Forgery Be Detected by the Same Model? Forensim: An Attention-Based State-Space Approach

  • 用注意力机制捕捉图像内部相似性,识别复制痕迹。
  • 在标准数据集上达到当前最佳性能,支持三类掩码输出。
  • 适合媒体审核、司法取证等需要精确溯源的场景。

我们提出 Forensim,一种基于注意力机制的状态空间框架,用于联合定位图像篡改(目标)和来源区域。不同于仅依赖伪影线索的传统方法,Forensim 能捕捉理解上下文所必需的复制模式。例如,在抗议图像中,仅检测出被复制的暴力行为而忽略其来源,可能误导解读,凸显联合定位的重要性。Forensim 输出三类掩码(原始、来源、目标),统一架构下支持检测拼接与复制-粘贴伪造。我们设计了一种视觉状态空间模型,利用归一化注意力图识别内部相似性,并结合基于区域的块注意力模块区分篡改区域。该设计支持端到端训练与精准定位。Forensim 在标准基准上表现优异。我们还发布了新数据集 CMFD-Anything,弥补现有复制-粘贴伪造数据集的不足。

原文摘要 · Abstract (English)

We introduce Forensim, an attention-based state-space framework for image forgery detection that jointly localizes both manipulated (target) and source regions. Unlike traditional approaches that rely solely on artifact cues to detect spliced or forged areas, Forensim is designed to capture duplication patterns crucial for understanding context. In scenarios such as protest imagery, detecting only the forged region, for example a duplicated act of violence inserted into a peaceful crowd, can mislead interpretation, highlighting the need for joint source-target localization. Forensim outputs three-class masks (pristine, source, target) and supports detection of both splicing and copy-move forgeries within a unified architecture. We propose a visual state-space model that leverages normalized attention maps to identify internal similarities, paired with a region-based block attention module to distinguish manipulated regions. This design enables end-to-end training and precise localization. Forensim achieves state-of-the-art performance on standard benchmarks. We also release CMFD-Anything, a new dataset addressing limitations of existing copy-move forgery datasets.

图像伪造注意力机制溯源检测统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。