arXiv:2506.20922cs.CV2025-06ICCV被引 11

M2SFormer通过多尺度注意力和边缘感知难度引导,精准定位图像伪造区域。

M2SFormer: Multi-Spectral and Multi-Scale Attention with Edge-Aware Difficulty Guidance for Image Forgery Localization

  • 融合多频谱与多尺度注意力,统一处理空间与频率特征。
  • 在多个数据集上优于现有模型,对未见域泛化能力更强。
  • 引入曲率难易度图指导细节保留,适合检测细微篡改。

图像编辑技术快速发展,既推动创新应用也加剧恶意篡改风险。基于深度学习的像素级伪造定位方法虽精度高,但常面临计算开销大、表征能力弱的问题,尤其对细微或复杂的篡改效果不佳。本文提出M2SFormer,一种基于Transformer编码器的新框架,突破传统分步处理空间与频率信息的局限,将多频谱与多尺度注意力统一于跳跃连接中,利用全局上下文更有效捕捉多样伪造痕迹。同时,为缓解上采样过程中的细节丢失,引入全局先验图与曲率度量(表示伪造定位难度),驱动难度感知注意力模块,增强对细微篡改的保留能力。在多个基准数据集上的大量实验表明,M2SFormer超越现有最先进模型,在跨域未见场景中表现出更优的检测与定位性能。

原文摘要 · Abstract (English)

Image editing techniques have rapidly advanced, facilitating both innovative use cases and malicious manipulation of digital images. Deep learning-based methods have recently achieved high accuracy in pixel-level forgery localization, yet they frequently struggle with computational overhead and limited representation power, particularly for subtle or complex tampering. In this paper, we propose M2SFormer, a novel Transformer encoder-based framework designed to overcome these challenges. Unlike approaches that process spatial and frequency cues separately, M2SFormer unifies multi-frequency and multi-scale attentions in the skip connection, harnessing global context to better capture diverse forgery artifacts. Additionally, our framework addresses the loss of fine detail during upsampling by utilizing a global prior map, a curvature metric indicating the difficulty of forgery localization, which then guides a difficulty-guided attention module to preserve subtle manipulations more effectively. Extensive experiments on multiple benchmark datasets demonstrate that M2SFormer outperforms existing state-of-the-art models, offering superior generalization in detecting and localizing forgeries across unseen domains.

图像伪造Transformer注意力机制细粒度检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。