用场景图融合多模态信息,提升假消息真实性判断准确率
SceneGraMMi: Scene Graph-boosted Hybrid-fusion for Multi-Modal Misinformation Veracity Prediction
- 通过跨模态场景图融合视觉与文本语义
- 在四个数据集上均超越现有最优方法
- 支持可解释性分析,适合内容安全研究者
虚假信息损害个人认知并影响社会叙事。尽管多模态假消息检测受到广泛关注,现有方法仍难以有效捕捉语义线索、关键区域及跨模态相似性。本文提出 SceneGraMMi——一种基于场景图增强的混合融合多模态假消息真实性预测方法,通过在不同模态间整合场景图以提升检测性能。在四个基准数据集上的实验表明,SceneGraMMi 均持续优于当前最先进方法。全面的消融实验验证了各组件贡献,同时采用 Shapley 值分析模型决策过程的可解释性。
原文摘要 · Abstract (English)
Misinformation undermines individual knowledge and affects broader societal narratives. Despite growing interest in the research community in multi-modal misinformation detection, existing methods exhibit limitations in capturing semantic cues, key regions, and cross-modal similarities within multi-modal datasets. We propose SceneGraMMi, a Scene Graph-boosted Hybrid-fusion approach for Multi-modal Misinformation veracity prediction, which integrates scene graphs across different modalities to improve detection performance. Experimental results across four benchmark datasets show that SceneGraMMi consistently outperforms state-of-the-art methods. In a comprehensive ablation study, we highlight the contribution of each component, while Shapley values are employed to examine the explainability of the model's decision-making process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。