arXiv:2507.01984cs.LGcs.CL2025-07被引 3

融合文本、图像与社交特征,提升假信息识别准确率。

Multimodal Misinformation Detection Using Early Fusion of Linguistic, Visual, and Social Features

  • 早期融合文本、图像与社交特征构建分类模型。
  • 多模态融合使准确率比单模态高15%,比双模态高5%。
  • 适用于选举与疫情等危机场景的假信息监测。

在选举和危机期间,社交媒体上充斥着大量虚假信息,已有研究主要集中在基于文本或图像的检测方法。然而,仅有少数研究探索了多模态特征的组合,例如将文本与图像结合以构建假信息检测模型。本研究通过早期融合策略,整合文本、图像及社交特征,分析了1,529条来自推特(现为X)的含图文微博,涵盖新冠疫情期间和选举阶段的数据。通过目标检测和光学字符识别(OCR)等技术,对数据进行增强,提取视觉与社交特征。结果表明,结合无监督与有监督机器学习模型,可使分类性能相比单模态模型提升15%,相比双模态模型提升5%。此外,研究还基于假信息微博特征及其传播者属性,分析了其传播模式。

原文摘要 · Abstract (English)

Amid a tidal wave of misinformation flooding social media during elections and crises, extensive research has been conducted on misinformation detection, primarily focusing on text-based or image-based approaches. However, only a few studies have explored multimodal feature combinations, such as integrating text and images for building a classification model to detect misinformation. This study investigates the effectiveness of different multimodal feature combinations, incorporating text, images, and social features using an early fusion approach for the classification model. This study analyzed 1,529 tweets containing both text and images during the COVID-19 pandemic and election periods collected from Twitter (now X). A data enrichment process was applied to extract additional social features, as well as visual features, through techniques such as object detection and optical character recognition (OCR). The results show that combining unsupervised and supervised machine learning models improves classification performance by 15% compared to unimodal models and by 5% compared to bimodal models. Additionally, the study analyzes the propagation patterns of misinformation based on the characteristics of misinformation tweets and the users who disseminate them.

假信息检测多模态融合社交网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。