arXiv:2608.08009cs.CVcs.AI2026-08中稿 · ACM MM 2026

提出可追溯证据的多模态伪造检测框架,让模型解释更可信。

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation

论文配图:Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation
图 1 · 摘自论文原文
  • 采用锚定-验证推理链,先分模态感知再跨模态比对。
  • 在多个数据集上达到顶尖检测性能,且解释与证据空间对齐。
  • 适合需要透明推理的司法鉴定与内容审核场景。

虚假新闻越来越多依赖跨模态图像-文本伪造,亟需可透明、可验证的推理链条来实现多模态媒体伪造检测与定位(DGM4)。现有方法生成黑箱检测结果,缺乏决策依据,限制了其在司法取证中的可靠性。多模态大语言模型(MLLMs)为可解释性提供了自然路径,但应用于DGM4面临两大挑战:一是模型生成的解释常与预测证据位置脱节,导致无验证归因;二是保证证据-结论一致性需主动优化,但统一训练信号难以区分定位与分类令牌,使多头联合训练不可靠。本文提出基于证据锚定取证推理(EFR)框架的多模态伪造检测器。EFR引入锚定-验证推理链,在跨模态对比前进行模态隔离感知,并以结论坐标作为显式锚点,要求下游证据必须在空间上对应。通过可验证奖励机制在训练中强制证据-结论一致性,结合模态解耦优势路由(MDA)机制缓解任务间信用错配。实验表明,EFR在多个基准数据集上达到当前最优性能,同时生成结构化取证推理记录,明确关联解释与证据位置。

原文摘要 · Abstract (English)

Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4). Existing methods produce black-box detection results without any decision rationale, limiting their reliability in forensic practice. Multi-modal Large Language Models (MLLMs) offer a natural path toward explainability, but applying them to DGM4 raises two difficulties. First, models tend to generate explanations disconnected from predicted evidence locations, producing unverified attribution. Second, enforcing evidence-conclusion consistency requires active optimization, yet uniform training signals fail to distinguish localization tokens from classification tokens, making multi-head joint training unreliable. We propose a multi-modal manipulation detector based on an Evidence-Grounded Forensic Reasoning (EFR) framework. EFR introduces an Anchor-and-Verify reasoning chain that enforces modality-isolated perception before cross-modal comparison, with conclusion coordinates as explicit anchors to which downstream evidence must spatially correspond. A verifiable reward system then enforces evidence-conclusion consistency during training, while a Modality-Decoupled Advantage (MDA) routing mechanism mitigats credit misassignment across prediction tasks. Experiments show that EFR achieves state-of-the-art performance while producing structured forensic reasoning records that explicitly bind explanations to evidence.

多模态检测可解释性证据锚定伪造识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。