多模态事实核查框架,能高效识别图文混杂的假信息
MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification
- 通过元素级差值与乘积融合跨模态证据,实现透明交互
- 在Factify-2和Mocheg上性能显著提升,抗噪声能力强
- 轻量高效适合记者、社会组织等实际场景部署
网络虚假信息日益以图文结合、包含OCR内容的多模态形式出现,严重损害公众信任与弱势群体权益。现有事实核查系统多依赖单模态信号或浅层融合,难以应对跨模态隐性矛盾。我们提出MultiCheck,一种轻量且可解释的多模态事实核查框架,联合分析文本、图像和OCR证据。其核心为基于元素级差值与乘积操作的关联融合模块,实现显式跨模态交互,计算开销极小。对比对齐目标进一步帮助模型区分支持与反驳证据,同时保持低内存与能耗,适合低资源环境部署。在Factify-2(5类)和Mocheg(3类)基准测试中,MultiCheck表现优异,即使在存在噪声OCR或缺失模态时仍保持鲁棒性。其高效性、可解释性与现实适应性,使其适用于新闻从业者、民间组织及网络可信度建设者,助力构建更安全可信的网络空间。
原文摘要 · Abstract (English)
Misinformation on the web increasingly appears in multimodal forms, combining text, images, and OCR-rendered content in ways that amplify harm to public trust and vulnerable communities. While prior fact-checking systems often rely on unimodal signals or shallow fusion strategies, modern misinformation campaigns operate across modalities and require models that can reason over subtle cross-modal inconsistencies in a transparent and responsible manner. We introduce MultiCheck, a lightweight and interpretable framework for multimodal fact verification that jointly analyzes textual, visual, and OCR evidence. At its core, MultiCheck employs a relational fusion module based on element-wise difference and product operations, allowing for explicit cross-modal interaction modeling with minimal computational overhead. A contrastive alignment objective further helps the model distinguish between supporting and refuting evidence while maintaining a small memory and energy footprint, making it suitable for low-resource deployment. Evaluated on the Factify-2 (5-class) and Mocheg (3-class) benchmarks, MultiCheck achieves huge performance improvement and remains robust under noisy OCR and missing modality conditions. Its efficiency, transparency, and real-world robustness make it well-suited for journalists, civil society organisations, and web integrity efforts working to build a safer and more trustworthy web.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。