arXiv:2506.19635cs.CRcs.AI2025-06被引 23

用简单特征检测新僵尸账号,效果竟不输复杂模型。

On the efficacy of old features for the detection of new bots

  • 用账号基础信息和客户端数据做分类,成本低且高效
  • 在6个新数据集上,简单特征表现接近顶尖检测器
  • 适合平台快速部署,尤其对资源有限的团队

过去十余年,学术界和平台方一直致力于解决僵尸账号检测问题。恶意机器人被用于散布垃圾信息、操纵公众人物,最终扭曲舆论。为应对这一挑战,已有多种方法被提出,主要基于各类分类器,利用从推特公开API获取的数据,提取从简单到复杂的账户特征。本研究以推特为基准,对比四种先进特征集在识别新型僵尸账号上的表现:Botometer的输出评分(基于1000+特征)、账户资料与时间线特征、以及用户发推所用客户端信息。分析基于六个近期发布的推特账户数据集进行,结果表明,通用分类器结合低成本计算的账户特征,已具备有效检测演化型僵尸账号的能力。

原文摘要 · Abstract (English)

For more than a decade now, academicians and online platform administrators have been studying solutions to the problem of bot detection. Bots are computer algorithms whose use is far from being benign: malicious bots are purposely created to distribute spam, sponsor public characters and, ultimately, induce a bias within the public opinion. To fight the bot invasion on our online ecosystem, several approaches have been implemented, mostly based on (supervised and unsupervised) classifiers, which adopt the most varied account features, from the simplest to the most expensive ones to be extracted from the raw data obtainable through the Twitter public APIs. In this exploratory study, using Twitter as a benchmark, we compare the performances of four state-of-art feature sets in detecting novel bots: one of the output scores of the popular bot detector Botometer, which considers more than 1,000 features of an account to take a decision; two feature sets based on the account profile and timeline; and the information about the Twitter client from which the user tweets. The results of our analysis, conducted on six recently released datasets of Twitter accounts, hint at the possible use of general-purpose classifiers and cheap-to-compute account features for the detection of evolved bots.

僵尸账号特征检测推特数据轻量级模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。