用机器学习识别学术文章在推特上的机器人刷量,发现健康类研究最易被操纵。
Public interest in science or bots? Selective amplification of scientific articles on Twitter
- 结合多源数据训练模型,通过文章特征判断推特传播中是否存在过度机器人活动。
- 模型识别准确率达0.70,健康与人类科学类文章更易受机器人影响。
- 为检测学术传播中的虚假热度提供工具,适合关注科研舆情的研究者使用。
社交平台使学术论文能迅速触达公众,但机器人刷量可能扭曲公共讨论,影响现实政策。本研究结合Altmetric数据、推特API及Botometer API,构建包含学术文章、文章特征和机器人活动标签的大型数据集。分析发现,健康与人类科学类研究更易遭遇机器人异常传播。基于该数据集训练的机器学习模型可实现0.70的准确率,有效识别潜在机器人活动。研究虽未评估机器人行为的恶意性,但提供了检测学术传播中虚假热度的工具,并为后续研究奠定基础。
原文摘要 · Abstract (English)
With the remarkable capability to reach the public instantly, social media has become integral in sharing scholarly articles to measure public response. Since spamming by bots on social media can steer the conversation and present a false public interest in given research, affecting policies impacting the public's lives in the real world, this topic warrants critical study and attention. We used the Altmetric dataset in combination with data collected through the Twitter Application Programming Interface (API) and the Botometer API. We combined the data into an extensive dataset with academic articles, several features from the article and a label indicating whether the article had excessive bot activity on Twitter or not. We analyzed the data to see the possibility of bot activity based on different characteristics of the article. We also trained machine-learning models using this dataset to identify possible bot activity in any given article. Our machine-learning models were capable of identifying possible bot activity in any academic article with an accuracy of 0.70. We also found that articles related to "Health and Human Science" are more prone to bot activity compared to other research areas. Without arguing the maliciousness of the bot activity, our work presents a tool to identify the presence of bot activity in the dissemination of an academic article and creates a baseline for future research in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。