arXiv:2509.15578cs.CVcs.AI2025-09被引 1

融合音视频文本信息,提升短视频假新闻识别准确率

Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality Fusion

  • 动态调整多模态权重,适应数据缺失场景
  • 在两个数据集上分别提升2.71%和4.14%的F1分数
  • 适合关注短视频虚假信息检测的研究者与平台方

短视频平台的迅速发展对假新闻检测提出了更高要求。当前方法难以应对短视频内容动态、多模态的特性。本文提出HFN(异构融合网络),整合视频、音频与文本信息,通过决策网络动态调整各模态权重,并引入加权多模态特征融合模块,增强不完整数据下的鲁棒性。同时构建了专门用于短视频假新闻检测的公开数据集VESV(VEracity on Short Videos)。在FakeTT与新收集的VESV数据集上的实验表明,该方法相比现有最优模型在Marco F1上分别提升2.71%和4.14%。本工作为复杂短视频生态中的假新闻识别提供了可靠解决方案,推动了反虚假信息技术的发展。

原文摘要 · Abstract (English)

The rapid proliferation of short video platforms has necessitated advanced methods for detecting fake news. This need arises from the widespread influence and ease of sharing misinformation, which can lead to significant societal harm. Current methods often struggle with the dynamic and multimodal nature of short video content. This paper presents HFN, Heterogeneous Fusion Net, a novel multimodal framework that integrates video, audio, and text data to evaluate the authenticity of short video content. HFN introduces a Decision Network that dynamically adjusts modality weights during inference and a Weighted Multi-Modal Feature Fusion module to ensure robust performance even with incomplete data. Additionally, we contribute a comprehensive dataset VESV (VEracity on Short Videos) specifically designed for short video fake news detection. Experiments conducted on the FakeTT and newly collected VESV datasets demonstrate improvements of 2.71% and 4.14% in Marco F1 over state-of-the-art methods. This work establishes a robust solution capable of effectively identifying fake news in the complex landscape of short video platforms, paving the way for more reliable and comprehensive approaches in combating misinformation.

假新闻检测多模态学习短视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。