针对社交媒体谣言检测数据不平衡问题,提出基于图对比学习的异常检测框架。
Towards Real-World Rumor Detection: Anomaly Detection Framework with Graph Supervised Contrastive Learning
- 将未标注数据视为正常内容,用图对比学习捕捉谣言与常规内容差异。
- 在微博、推特构建大规模对话数据集,验证模型在少样本和数据不平衡下的优势。
- 适用于真实场景中谣言占比极低的社交媒体环境,适合反虚假信息研究者。
现有基于传播结构学习的谣言检测方法多将任务视为类别平衡的分类问题,但真实社交网络中谣言仅占极少数,数据严重不平衡。为此,我们从微博和推特构建两个大规模对话数据集,并分析领域分布差异:非谣言主要集中于娱乐类,而谣言则集中在新闻类,表明谣言检测本质上符合异常检测范式。据此,我们提出图监督对比学习的异常检测框架(AD-GSCL),启发式地将未标注数据视为正常内容,适配图对比学习用于谣言识别。大量实验表明,该方法在类别平衡、数据不平衡及小样本条件下均表现更优。研究结果为真实世界中数据分布不均衡的谣言检测提供了重要启示。
原文摘要 · Abstract (English)
Current rumor detection methods based on propagation structure learning predominately treat rumor detection as a class-balanced classification task on limited labeled data. However, real-world social media data exhibits an imbalanced distribution with a minority of rumors among massive regular posts. To address the data scarcity and imbalance issues, we construct two large-scale conversation datasets from Weibo and Twitter and analyze the domain distributions. We find obvious differences between rumor and non-rumor distributions, with non-rumors mostly in entertainment domains while rumors concentrate in news, indicating the conformity of rumor detection to an anomaly detection paradigm. Correspondingly, we propose the Anomaly Detection framework with Graph Supervised Contrastive Learning (AD-GSCL). It heuristically treats unlabeled data as non-rumors and adapts graph contrastive learning for rumor detection. Extensive experiments demonstrate AD-GSCL's superiority under class-balanced, imbalanced, and few-shot conditions. Our findings provide valuable insights for real-world rumor detection featuring imbalanced data distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。