检测图像像素伪造的模型对多模态假信息验证帮助有限,反而可能误导判断。
Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection
- 用多模态基准测试对比纯图像检测与证据驱动的推理系统
- 像素级检测器在两数据集上F1仅0.26-0.53,注入后反使性能下降0.04-0.08
- 真正有效的是结合外部证据和多智能体辩论的语义理解系统
在多模态虚假信息中,欺骗性不仅源于图像的像素级篡改,更来自图像与文本共同表达的语义和上下文主张。然而,多数深伪检测器仅关注像素级伪造,未考虑主张层面的意义,尽管它们正被越来越多地集成到自动化事实核查(AFC)流程中。这引发一个核心科学与实践问题:像素级检测器是否为验证图像-文本主张提供有用信号,还是引入误导性的真实性先验,破坏基于证据的推理?我们首次系统分析了深伪检测器在多模态虚假信息检测中的作用。通过两个互补的基准测试MMFakeBench和DGM4,评估了三种方法:(1)最先进的仅图像深伪检测器,(2)基于证据的事实核查系统,该系统通过蒙特卡洛树搜索(MCTS)引导检索,并通过多智能体辩论(MAD)进行推理性推理;(3)将检测器输出作为辅助证据的混合系统。结果显示,深伪检测器独立使用时表现有限,于MMFakeBench上F1为0.26–0.53,于DGM4上为0.33–0.49;而将其预测结果引入核查流程后,性能稳定下降0.04–0.08 F1,原因在于非因果的真实性假设。相比之下,以证据为中心的核查系统表现最优,在MMFakeBench上达到约0.81的F1,在DGM4上达0.55。总体而言,研究发现多模态主张验证主要依赖语义理解与外部证据,像素级伪造信号无法可靠提升现实世界图像-文本虚假信息的推理能力。
原文摘要 · Abstract (English)
In multimodal misinformation, deception usually arises not just from pixel-level manipulations in an image, but from the semantic and contextual claim jointly expressed by the image-text pair. Yet most deepfake detectors, engineered to detect pixel-level forgeries, do not account for claim-level meaning, despite their growing integration in automated fact-checking (AFC) pipelines. This raises a central scientific and practical question: Do pixel-level detectors contribute useful signal for verifying image-text claims, or do they instead introduce misleading authenticity priors that undermine evidence-based reasoning? We provide the first systematic analysis of deepfake detectors in the context of multimodal misinformation detection. Using two complementary benchmarks, MMFakeBench and DGM4, we evaluate: (1) state-of-the-art image-only deepfake detectors, (2) an evidence-driven fact-checking system that performs tool-guided retrieval via Monte Carlo Tree Search (MCTS) and engages in deliberative inference through Multi-Agent Debate (MAD), and (3) a hybrid fact-checking system that injects detector outputs as auxiliary evidence. Results across both benchmark datasets show that deepfake detectors offer limited standalone value, achieving F1 scores in the range of 0.26-0.53 on MMFakeBench and 0.33-0.49 on DGM4, and that incorporating their predictions into fact-checking pipelines consistently reduces performance by 0.04-0.08 F1 due to non-causal authenticity assumptions. In contrast, the evidence-centric fact-checking system achieves the highest performance, reaching F1 scores of approximately 0.81 on MMFakeBench and 0.55 on DGM4. Overall, our findings demonstrate that multimodal claim verification is driven primarily by semantic understanding and external evidence, and that pixel-level artifact signals do not reliably enhance reasoning over real-world image-text misinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。