arXiv:2601.02566cs.CV2026-01

用视觉Mamba与图神经网络定位深伪和浅伪图像篡改区域

Shallow- and Deep-fake Image Manipulation Localization Using Vision Mamba and Guided Graph Neural Network

  • 结合视觉Mamba提取篡改边界特征
  • 提出引导式图神经网络增强真假像素区分度
  • 在深伪与浅伪图像上均表现优于现有方法

图像篡改定位是一项关键研究任务,因伪造图像可能对社会多个方面造成重大影响。此类篡改可通过传统图像编辑工具(称作“浅伪”)或先进人工智能技术(“深伪”)生成。尽管已有大量研究聚焦于浅伪图像或深伪视频的篡改定位,但同时处理两类情况的方法仍较少。本文探索了使用深度学习网络定位浅伪与深伪图像篡改的可行性,并提出相应解决方案。为精确区分真实与篡改像素,我们利用视觉Mamba网络提取清晰描述篡改与未篡改区域边界的特征图。为进一步强化该区分,我们提出一种新型引导式图神经网络(G-GNN)模块,以放大篡改与真实像素之间的差异。评估结果表明,所提方法在推理准确率上优于其他主流方法。

原文摘要 · Abstract (English)

Image manipulation localization is a critical research task, given that forged images may have a significant societal impact of various aspects. Such image manipulations can be produced using traditional image editing tools (known as "shallowfakes") or advanced artificial intelligence techniques ("deepfakes"). While numerous studies have focused on image manipulation localization on either shallowfake images or deepfake videos, few approaches address both cases. In this paper, we explore the feasibility of using a deep learning network to localize manipulations in both shallow- and deep-fake images, and proposed a solution for such purpose. To precisely differentiate between authentic and manipulated pixels, we leverage the Vision Mamba network to extract feature maps that clearly describe the boundaries between tampered and untouched regions. To further enhance this separation, we propose a novel Guided Graph Neural Network (G-GNN) module that amplifies the distinction between manipulated and authentic pixels. Our evaluation results show that our proposed method achieved higher inference accuracy compared to other state-of-the-art methods.

图像伪造视觉Mamba图神经网络篡改定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。