arXiv:2508.04252cs.SIcs.CL2025-08被引 3

用海量无标签话题数据提升谣言检测模型泛化能力

Graph Representation Learning with Massive Unlabeled Data for Rumor Detection

  • 利用微博和推特的无标签传播结构数据进行图自监督学习
  • 在少样本条件下优于专用谣言检测模型,10年跨度数据验证有效性
  • 适合关注跨事件泛化、少样本谣言检测的研究者

随着社交媒体发展,谣言传播迅速,对社会经济造成重大危害。现有基于谣言传播结构的学习方法虽有效,但受限于大规模标注数据难以获取,导致模型泛化能力差,在新事件上性能下降。为此,本研究从微博和推特爬取包含声明传播结构的大规模无标签话题数据,用于提升图表示学习模型的语义理解能力。采用InfoGraph、JOAO和GraphMAE三种典型图自监督方法,结合两种常用训练策略,验证通用图半监督方法在谣言检测任务中的表现。此外,为缓解无标签数据与谣言数据间的时间与话题差异,我们还收集了覆盖十年(2022年前)多种话题的微博辟谣平台谣言数据集。实验表明,这些通用图自监督学习方法在少样本条件下表现优于以往专用于谣言检测的方法,证明了借助大规模无标签数据可显著提升模型泛化能力。

原文摘要 · Abstract (English)

With the development of social media, rumors spread quickly, cause great harm to society and economy. Thereby, many effective rumor detection methods have been developed, among which the rumor propagation structure learning based methods are particularly effective compared to other methods. However, the existing methods still suffer from many issues including the difficulty to obtain large-scale labeled rumor datasets, which leads to the low generalization ability and the performance degeneration on new events since rumors are time-critical and usually appear with hot topics or newly emergent events. In order to solve the above problems, in this study, we used large-scale unlabeled topic datasets crawled from the social media platform Weibo and Twitter with claim propagation structure to improve the semantic learning ability of a graph reprentation learing model on various topics. We use three typical graph self-supervised methods, InfoGraph, JOAO and GraphMAE in two commonly used training strategies, to verify the performance of general graph semi-supervised methods in rumor detection tasks. In addition, for alleviating the time and topic difference between unlabeled topic data and rumor data, we also collected a rumor dataset covering a variety of topics over a decade (10-year ago from 2022) from the Weibo rumor-refuting platform. Our experiments show that these general graph self-supervised learning methods outperform previous methods specifically designed for rumor detection tasks and achieve good performance under few-shot conditions, demonstrating the better generalization ability with the help of our massive unlabeled topic dataset.

谣言检测图神经网络自监督学习少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。