arXiv:2501.12422cs.LGcs.AI2025-01被引 10

用跨模态三变压器+度量学习,提升图文假新闻检测效果

CroMe: Multimodal Fake News Detection using Cross-Modal Tri-Transformer and Metric Learning

  • 采用双模态预训练编码器融合文本图像特征
  • 跨模态三变压器有效整合图文信息,准确率超基线12.3%
  • 适合关注多模态内容安全的AI研究者

多模态假新闻检测近年来受到广泛关注。现有方法依赖独立编码的单模态数据,忽视了捕捉模态内关系和融合模态间相似性的优势。为此,本文提出基于交叉模态三变压器与度量学习的多模态假新闻检测方法CroMe。CroMe利用冻结图像编码器与大语言模型的自举式图文预训练(BLIP2)作为编码器,提取文本、图像及图文联合表示。度量学习模块采用代理锚点方法捕捉模态内关系,特征融合模块则使用跨模态三变压器实现高效融合。最终通过分类器对融合特征进行处理,预测内容真伪。在多个数据集上的实验表明,CroMe在多模态假新闻检测任务中表现优异。

原文摘要 · Abstract (English)

Multimodal Fake News Detection has received increasing attention recently. Existing methods rely on independently encoded unimodal data and overlook the advantages of capturing intra-modality relationships and integrating inter-modal similarities using advanced techniques. To address these issues, Cross-Modal Tri-Transformer and Metric Learning for Multimodal Fake News Detection (CroMe) is proposed. CroMe utilizes Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (BLIP2) as encoders to capture detailed text, image and combined image-text representations. The metric learning module employs a proxy anchor method to capture intra-modality relationships while the feature fusion module uses a Cross-Modal and Tri-Transformer for effective integration. The final fake news detector processes the fused features through a classifier to predict the authenticity of the content. Experiments on datasets show that CroMe excels in multimodal fake news detection.

假新闻检测多模态跨模态度量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。