通过跨模态实体一致性检测视频谣言,提升识别准确率。
Detecting Misinformation in Multimedia Content through Cross-Modal Entity Consistency: A Dual Learning Approach
- 利用双学习框架建模多模态实体一致性
- 在多个数据集上超越现有最佳模型表现
- 适合关注视频内容安全与可信度研究者
社交媒体内容已从文本扩展至多模态形式,这给虚假信息的识别带来了新挑战。以往研究多集中于单一模态或图文组合,缺乏对多模态虚假信息的有效检测方法。尽管实体一致性概念在多模态虚假信息检测中具有潜力,但将高维表示简化为标量值会忽略不同模态间复杂性的差异。为此,本文提出一种多媒体虚假信息检测(MultiMD)框架,通过跨模态实体一致性分析视频内容中的虚假信息。所提出的双学习方法不仅提升了虚假信息检测性能,还优化了多模态实体一致性的表征学习。实验结果表明,MultiMD 在多个基准测试中优于当前最优模型,并验证了各模态在检测中的关键作用。本研究为多模态虚假信息检测提供了新的方法论和技术洞见。
原文摘要 · Abstract (English)
The landscape of social media content has evolved significantly, extending from text to multimodal formats. This evolution presents a significant challenge in combating misinformation. Previous research has primarily focused on single modalities or text-image combinations, leaving a gap in detecting multimodal misinformation. While the concept of entity consistency holds promise in detecting multimodal misinformation, simplifying the representation to a scalar value overlooks the inherent complexities of high-dimensional representations across different modalities. To address these limitations, we propose a Multimedia Misinformation Detection (MultiMD) framework for detecting misinformation from video content by leveraging cross-modal entity consistency. The proposed dual learning approach allows for not only enhancing misinformation detection performance but also improving representation learning of entity consistency across different modalities. Our results demonstrate that MultiMD outperforms state-of-the-art baseline models and underscore the importance of each modality in misinformation detection. Our research provides novel methodological and technical insights into multimodal misinformation detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。