无需标注数据,用图文对比学习+大模型识别假新闻。
A Self-Learning Multimodal Approach for Fake News Detection

- 用对比学习提取图文特征,不依赖人工标注。
- 在公开数据集上准确率、精确率、召回率、F1均超85%。
- 适合研究虚假信息检测与多模态学习的学者使用。
社交媒体的快速发展导致在线新闻内容激增,虚假信息传播日益严重。尽管机器学习已被广泛用于假新闻检测,但标注数据稀缺仍是关键挑战。虚假信息常以图文配对形式出现,即新闻标题或文章伴随相关图像。本文提出一种自学习多模态假新闻分类模型。该模型利用对比学习进行特征提取,无需标注数据;同时融合大型语言模型(LLMs)的优势,联合分析文本与图像特征。得益于LLMs在海量语料训练下的语言处理能力,模型表现出色。在公开数据集上的实验表明,所提方法优于多个前沿分类模型,各项指标(准确率、精确率、召回率、F1分数)均超过85%,验证了其在多模态假新闻检测中的有效性。
原文摘要 · Abstract (English)
The rapid growth of social media has resulted in an explosion of online news content, leading to a significant increase in the spread of misleading or false information. While machine learning techniques have been widely applied to detect fake news, the scarcity of labeled datasets remains a critical challenge. Misinformation frequently appears as paired text and images, where a news article or headline is accompanied by a related visuals. In this paper, we introduce a self-learning multimodal model for fake news classification. The model leverages contrastive learning, a robust method for feature extraction that operates without requiring labeled data, and integrates the strengths of Large Language Models (LLMs) to jointly analyze both text and image features. LLMs are excel at this task due to their ability to process diverse linguistic data drawn from extensive training corpora. Our experimental results on a public dataset demonstrate that the proposed model outperforms several state-of-the-art classification approaches, achieving over 85% accuracy, precision, recall, and F1-score. These findings highlight the model's effectiveness in tackling the challenges of multimodal fake news detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。