通过因果干预消除多模态假新闻检测中的混淆因素,提升模型可靠性。
Deconfounded Reasoning for Multimodal Fake News Detection via Causal Intervention
- 构建统一因果模型,识别三类跨模态混淆因子
- 在FakeSV和FVC数据集上分别提升4.27%和4.80%准确率
- 适合关注假新闻检测鲁棒性与可解释性的研究者
社交媒体的快速发展导致文本、图像、音频和视频等多种形式的假新闻广泛传播。传统单模态检测方法难以应对复杂的跨模态篡改,因此多模态假新闻检测成为更有效的解决方案。然而,现有方法常忽略复杂跨模态交互中隐藏的混淆因素,导致模型依赖虚假统计相关性而非真实因果机制。本文提出基于因果干预的多模态去混淆检测框架(CIMDD),通过统一的结构因果模型(SCM)系统建模三类混淆因子:词汇语义混淆因子(LSC)、潜在视觉混淆因子(LVC)和动态跨模态耦合混淆因子(DCCC)。为缓解其影响,设计三种因果模块,分别基于后门调整、前门调整和跨模态联合干预,从不同角度阻断虚假相关性,实现表示的因果解耦。在FakeSV和FVC数据集上的实验表明,CIMDD显著提升检测准确率,分别优于当前最优方法4.27%和4.80%。大量实验还显示,该框架在多种多模态场景下具备强泛化性和鲁棒性。
原文摘要 · Abstract (English)
The rapid growth of social media has led to the widespread dissemination of fake news across multiple content forms, including text, images, audio, and video. Traditional unimodal detection methods fall short in addressing complex cross-modal manipulations; as a result, multimodal fake news detection has emerged as a more effective solution. However, existing multimodal approaches, especially in the context of fake news detection on social media, often overlook the confounders hidden within complex cross-modal interactions, leading models to rely on spurious statistical correlations rather than genuine causal mechanisms. In this paper, we propose the Causal Intervention-based Multimodal Deconfounded Detection (CIMDD) framework, which systematically models three types of confounders via a unified Structural Causal Model (SCM): (1) Lexical Semantic Confounder (LSC); (2) Latent Visual Confounder (LVC); (3) Dynamic Cross-Modal Coupling Confounder (DCCC). To mitigate the influence of these confounders, we specifically design three causal modules based on backdoor adjustment, frontdoor adjustment, and cross-modal joint intervention to block spurious correlations from different perspectives and achieve causal disentanglement of representations for deconfounded reasoning. Experimental results on the FakeSV and FVC datasets demonstrate that CIMDD significantly improves detection accuracy, outperforming state-of-the-art methods by 4.27% and 4.80%, respectively. Furthermore, extensive experimental results indicate that CIMDD exhibits strong generalization and robustness across diverse multimodal scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。