提出首个语义对齐的多模态篡改检测方法,更贴近真实伪造场景。
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
- 构建首个语义一致的多模态篡改数据集SAMM,视觉与文本同步造假。
- 提出RamDG框架,结合外部知识检索提升检测精度,比现有方法高2.06%。
- 适合媒体取证、AI安全研究者,尤其关注真实场景下内容伪造检测。
多模态数据中的篡改内容检测已成为媒体取证的关键挑战。现有基准虽展示技术进步,但存在误对齐缺陷:实际攻击通常保持跨模态语义一致性,而当前数据集人为破坏跨模态对齐,产生易被识别的异常。为弥合这一差距,我们首次提出检测语义协调的多模态篡改,即视觉修改与语义一致的文本描述系统性匹配。我们构建首个语义对齐多模态篡改(SAMM)数据集,采用两阶段流程:1)应用前沿图像篡改技术;2)生成上下文合理的文本叙事以强化视觉欺骗。基于此,我们提出检索增强的篡改检测与定位框架(RamDG)。RamDG首先利用外部知识库检索上下文证据,作为辅助文本,与输入共同编码,通过图像伪造定位和深度篡改检测模块追踪所有篡改痕迹。大量实验表明,该框架在SAMM数据集上检测准确率比现有最优方法高出2.06%。数据集与代码已公开于https://github.com/shen8424/SAMM-RamDG-CAP。
原文摘要 · Abstract (English)
The detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifacts that poorly reflect real-world manipulation patterns: practical attacks typically maintain semantic consistency across modalities, whereas current datasets artificially disrupt cross-modal alignment, creating easily detectable anomalies. To bridge this gap, we pioneer the detection of semantically-coordinated manipulations where visual edits are systematically paired with semantically consistent textual descriptions. Our approach begins with constructing the first Semantic-Aligned Multimodal Manipulation (SAMM) dataset, generated through a two-stage pipeline: 1) applying state-of-the-art image manipulations, followed by 2) generation of contextually-plausible textual narratives that reinforce the visual deception. Building on this foundation, we propose a Retrieval-Augmented Manipulation Detection and Grounding (RamDG) framework. RamDG commences by harnessing external knowledge repositories to retrieve contextual evidence, which serves as the auxiliary texts and encoded together with the inputs through our image forgery grounding and deep manipulation detection modules to trace all manipulations. Extensive experiments demonstrate our framework significantly outperforms existing methods, achieving 2.06\% higher detection accuracy on SAMM compared to state-of-the-art approaches. The dataset and code are publicly available at https://github.com/shen8424/SAMM-RamDG-CAP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。