arXiv:2505.23449cs.MMcs.CV2025-05ACL被引 10

用深层语义关联提升图文错位谣言检测准确率

CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection

  • 通过共现关系生成识别图文间接语义关联
  • 在多个数据集上比现有方法提升5.2%以上准确率
  • 适合需要可解释性的谣言检测场景

多模态大语言模型(MLLM)在视觉推理和文本生成方面表现出色。尽管已有研究尝试用MLLM检测图文错位(OOC)谣言,但我们的实证分析揭示了两个持续存在的挑战:在直接推理与证据增强推理下评估GPT-4o模型,结果显示其难以捕捉深层关系——即图像与文本无直接关联,但存在潜在语义联系的案例;此外,证据中的噪声进一步降低了检测精度。为此,我们提出CMIE框架,包含共现关系生成(CRG)策略与关联评分(AS)机制,能识别图文间的潜在共存关系,并选择性利用相关证据以提升检测效果。实验表明,该方法显著优于现有方法。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have demonstrated impressive capabilities in visual reasoning and text generation. While previous studies have explored the application of MLLM for detecting out-of-context (OOC) misinformation, our empirical analysis reveals two persisting challenges of this paradigm. Evaluating the representative GPT-4o model on direct reasoning and evidence augmented reasoning, results indicate that MLLM struggle to capture the deeper relationships-specifically, cases in which the image and text are not directly connected but are associated through underlying semantic links. Moreover, noise in the evidence further impairs detection accuracy. To address these challenges, we propose CMIE, a novel OOC misinformation detection framework that incorporates a Coexistence Relationship Generation (CRG) strategy and an Association Scoring (AS) mechanism. CMIE identifies the underlying coexistence relationships between images and text, and selectively utilizes relevant evidence to enhance misinformation detection. Experimental results demonstrate that our approach outperforms existing methods.

谣言检测多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。