arXiv:2504.19212cs.CVcs.AI2025-04被引 2

用多模态胶囊网络检测指令引导的深度伪造图像

CapsFake: A Multimodal Capsule Network for Detecting Instruction-Guided Deepfakes

论文配图:CapsFake: A Multimodal Capsule Network for Detecting Instruction-Guided Deepfakes
图 1 · 摘自论文原文
  • 融合视觉、文本和频域特征的胶囊网络,动态定位篡改区域
  • 在多个数据集上准确率比现有方法最高提升20%
  • 对自然扰动和对抗攻击均保持94%以上检测率,适合真实场景

指令引导的图像编辑技术快速发展,通过真实图像与文本提示生成难以察觉的细微篡改,严重威胁数字图像真实性。现有检测手段对此类操作效果有限。本文提出多模态胶囊网络CapsFake,整合视觉、文本与频域低层胶囊,通过竞争性路由机制生成高层胶囊,精准识别篡改区域。在MagicBrush、Unsplash Edits、Open Images Edits及Multi-turn Edits等多样化数据集上,其检测准确率相比最先进方法最高提升20%。消融实验表明,在自然扰动下检测率超94%,对抗攻击下达96%,且具备优异的未见编辑场景泛化能力。该方法为应对复杂图像伪造提供了有效框架。

原文摘要 · Abstract (English)

The rapid evolution of deepfake technology, particularly in instruction-guided image editing, threatens the integrity of digital images by enabling subtle, context-aware manipulations. Generated conditionally from real images and textual prompts, these edits are often imperceptible to both humans and existing detection systems, revealing significant limitations in current defenses. We propose a novel multimodal capsule network, CapsFake, designed to detect such deepfake image edits by integrating low-level capsules from visual, textual, and frequency-domain modalities. High-level capsules, predicted through a competitive routing mechanism, dynamically aggregate local features to identify manipulated regions with precision. Evaluated on diverse datasets, including MagicBrush, Unsplash Edits, Open Images Edits, and Multi-turn Edits, CapsFake outperforms state-of-the-art methods by up to 20% in detection accuracy. Ablation studies validate its robustness, achieving detection rates above 94% under natural perturbations and 96% against adversarial attacks, with excellent generalization to unseen editing scenarios. This approach establishes a powerful framework for countering sophisticated image manipulations.

深度伪造多模态胶囊网络图像检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。