构建大规模社交网络图谱,区分真实趋势与协同操控的假热门
Large Engagement Networks for Classifying Coordinated Campaigns and Organic Twitter Trends
- 通过分析用户互动构建微博趋势的关联网络
- 数据集含179个操控事件和135个真实趋势,平均每个网络超万节点
- 为图分类模型提供挑战性测试场,适合研究虚假信息检测
社交媒体用户与非真实账号(如机器人)可能协同推广特定话题,制造出看似自发流行的假象。由于缺乏可靠的标注数据,判断话题是自然流行还是受控操控极具挑战。本文通过识别临时性操控攻击(即短期集中发帖后立即删除)来构建真实标签。我们手动整理了真实趋势数据集,并据此构建用户为节点、互动行为(转发、回复、引用)为边的参与网络。发布包含179个操控事件和135个非操控事件的大型参与网络数据集(LEN),每张图平均约1.1万个节点和2.3万条边。相比传统小规模图分类数据集,该数据集规模大得多。实验表明,现有先进图神经网络在辨别操控与非操控趋势及类型上表现一般。该数据集为大规模图分类提供了独特且具挑战性的研究平台,有助于推动图学习技术发展,并应用于识别协同操纵与真实趋势。
原文摘要 · Abstract (English)
Social media users and inauthentic accounts, such as bots, may coordinate in promoting their topics. Such topics may give the impression that they are organically popular among the public, even though they are astroturfing campaigns that are centrally managed. It is challenging to predict if a topic is organic or a coordinated campaign due to the lack of reliable ground truth. In this paper, we create such ground truth by detecting the campaigns promoted by ephemeral astroturfing attacks. These attacks push any topic to Twitter's (X) trends list by employing bots that tweet in a coordinated manner in a short period and then immediately delete their tweets. We manually curate a dataset of organic Twitter trends. We then create engagement networks out of these datasets which can serve as a challenging testbed for graph classification task to distinguish between campaigns and organic trends. Engagement networks consist of users as nodes and engagements as edges (retweets, replies, and quotes) between users. We release the engagement networks for 179 campaigns and 135 non-campaigns, and also provide finer-grain labels to characterize the type of the campaigns and non-campaigns. Our dataset, LEN (Large Engagement Networks), is available in the URL below. In comparison to traditional graph classification datasets, which are small with tens of nodes and hundreds of edges at most, graphs in LEN are larger. The average graph in LEN has ~11K nodes and ~23K edges. We show that state-of-the-art GNN methods give only mediocre results for campaign vs. non-campaign and campaign type classification on LEN. LEN offers a unique and challenging playfield for the graph classification problem. We believe that LEN will help advance the frontiers of graph classification techniques on large networks and also provide an interesting use case in terms of distinguishing coordinated campaigns and organic trends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。