arXiv:2505.23586cs.CVcs.MM2025-05被引 1

无需像素标注,用多尺度特征定位图像篡改区域。

Weakly-supervised Localization of Manipulated Image Regions Using Multi-resolution Learned Features

  • 利用图像级检测网络生成激活图,融合多尺度特征粗定位篡改区。
  • 结合预训练分割模型(如DeepLab)的区域信息,通过贝叶斯推理精修定位。
  • 在无像素标签场景下实现高精度篡改区域定位,适合真实世界应用。

数字图像爆炸式增长与图像编辑工具的普及,使图像篡改检测日益重要。现有深度学习方法虽在图像级分类上表现优异,但在篡改区域定位和可解释性方面不足。真实场景中缺乏像素级标注,限制了全监督定位技术的应用。为此,我们提出一种新型弱监督方法:将图像级篡改检测网络(如WCBnet)生成的激活图与预训练分割模型(如DeepLab、SegmentAnything、PSPnet)的分割图融合。首先生成多视图特征图进行粗略定位,再利用预训练模型提供的细粒度区域信息,通过贝叶斯推理优化定位结果。实验表明,该方法在不依赖像素级标签的前提下,有效实现了篡改区域定位。

原文摘要 · Abstract (English)

The explosive growth of digital images and the widespread availability of image editing tools have made image manipulation detection an increasingly critical challenge. Current deep learning-based manipulation detection methods excel in achieving high image-level classification accuracy, they often fall short in terms of interpretability and localization of manipulated regions. Additionally, the absence of pixel-wise annotations in real-world scenarios limits the existing fully-supervised manipulation localization techniques. To address these challenges, we propose a novel weakly-supervised approach that integrates activation maps generated by image-level manipulation detection networks with segmentation maps from pre-trained models. Specifically, we build on our previous image-level work named WCBnet to produce multi-view feature maps which are subsequently fused for coarse localization. These coarse maps are then refined using detailed segmented regional information provided by pre-trained segmentation models (such as DeepLab, SegmentAnything and PSPnet), with Bayesian inference employed to enhance the manipulation localization. Experimental results demonstrate the effectiveness of our approach, highlighting the feasibility to localize image manipulations without relying on pixel-level labels.

图像篡改弱监督定位分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。