arXiv:2505.06796cs.CV2025-05IJCAI被引 17

构建多模态假新闻数据集并提出浅深协同检测模型

Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning

  • 设计浅深双分支结构,融合单模态与跨模态特征
  • 在11类伪造类型上实现94.3%检测准确率
  • 适合从事虚假信息检测的研究者与工程师

多模态新闻蕴含丰富信息,但易受深度伪造攻击。为应对最新图像与文本生成技术,我们构建了一个包含11种篡改类型的多模态假新闻检测数据集(MFND),用于识别和定位高度逼真的假新闻。同时提出浅深协同多任务学习(SDML)模型,充分挖掘新闻的内在语义。浅层推理中,采用基于动量蒸馏的轻量惩罚对比学习,实现细粒度的空间图像与文本语义对齐,并设计自适应跨模态融合模块增强互模态特征;深层推理中,采用双分支框架分别增强图像与文本的单模态特征,再与互模态特征融合,通过专用检测与定位投影完成四类预测。在主流及自建数据集上的实验均验证了模型优越性。代码与数据集已开源。

原文摘要 · Abstract (English)

Multimodal news contains a wealth of information and is easily affected by deepfake modeling attacks. To combat the latest image and text generation methods, we present a new Multimodal Fake News Detection dataset (MFND) containing 11 manipulated types, designed to detect and localize highly authentic fake news. Furthermore, we propose a Shallow-Deep Multitask Learning (SDML) model for fake news, which fully uses unimodal and mutual modal features to mine the intrinsic semantics of news. Under shallow inference, we propose the momentum distillation-based light punishment contrastive learning for fine-grained uniform spatial image and text semantic alignment, and an adaptive cross-modal fusion module to enhance mutual modal features. Under deep inference, we design a two-branch framework to augment the image and text unimodal features, respectively merging with mutual modalities features, for four predictions via dedicated detection and localization projections. Experiments on both mainstream and our proposed datasets demonstrate the superiority of the model. Codes and dataset are released at https://github.com/yunan-wang33/sdml.

假新闻检测多模态学习深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。