arXiv:2508.05684cs.CRcs.LG2025-08被引 1

用动态融合机制提升多模态假新闻检测准确率

MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

  • 基于上下文自适应调整图文重要性权重
  • 在8万样本数据集上达到0.938的F1分数
  • 适合需要高精度假新闻识别的平台应用

社交媒体上多模态假新闻泛滥,严重威胁公众信任与社会稳定。传统以文本为主的检测方法常因文本与图像间的误导性配合而失效。尽管大视觉语言模型(LVLMs)为多模态理解带来希望,但如何有效融合信息量不均衡或矛盾的模态仍是关键挑战。本文提出MM-FusionNet,核心是上下文感知动态融合模块(CADFM),通过双向跨模态注意力和新颖的动态模态门控网络,自适应学习并分配文本与视觉特征的重要性权重,实现信息智能优先。在包含8万样本的大规模多模态假新闻数据集(LMFND)上,该模型取得0.938的F1-score,超越现有多模态基线约0.5%,显著优于单模态方法。进一步分析显示,模型具备动态加权能力、对模态扰动的鲁棒性,且性能接近人类水平,验证了其在真实场景中的有效性与可解释性。

原文摘要 · Abstract (English)

The proliferation of multi-modal fake news on social media poses a significant threat to public trust and social stability. Traditional detection methods, primarily text-based, often fall short due to the deceptive interplay between misleading text and images. While Large Vision-Language Models (LVLMs) offer promising avenues for multi-modal understanding, effectively fusing diverse modal information, especially when their importance is imbalanced or contradictory, remains a critical challenge. This paper introduces MM-FusionNet, an innovative framework leveraging LVLMs for robust multi-modal fake news detection. Our core contribution is the Context-Aware Dynamic Fusion Module (CADFM), which employs bi-directional cross-modal attention and a novel dynamic modal gating network. This mechanism adaptively learns and assigns importance weights to textual and visual features based on their contextual relevance, enabling intelligent prioritization of information. Evaluated on the large-scale Multi-modal Fake News Dataset (LMFND) comprising 80,000 samples, MM-FusionNet achieves a state-of-the-art F1-score of 0.938, surpassing existing multi-modal baselines by approximately 0.5% and significantly outperforming single-modal approaches. Further analysis demonstrates the model's dynamic weighting capabilities, its robustness to modality perturbations, and performance remarkably close to human-level, underscoring its practical efficacy and interpretability for real-world fake news detection.

假新闻检测多模态融合视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。