用元学习提升维基百科假账号检测精度,尤其擅长小数据下的行为模仿。
Detecting Sockpuppetry on Wikipedia Using Meta-Learning
- 通过元学习让模型快速适应新假账号群体的写作风格
- 在数据稀缺情况下,检测精度显著优于传统预训练模型
- 适合研究虚假账号识别与少样本学习的研究者
维基百科上的恶意假账号(sockpuppet)检测对维护网络信息可靠性至关重要,可防止虚假信息传播。以往机器学习方法依赖风格和元数据特征,但缺乏对作者个体行为的适应能力,尤其在文本数据有限时难以建模特定假账号群体。为此,我们提出应用元学习——一种在多任务中训练以提升小样本环境下性能的技术。该方法优化模型,使其能快速适应新假账号群体的写作风格。实验表明,相比预训练模型,元学习显著提升了预测精度,推动了开放编辑平台反假账号技术的发展。我们还发布了新的假账号调查数据集,以促进未来在假账号检测与元学习领域的研究。
原文摘要 · Abstract (English)
Malicious sockpuppet detection on Wikipedia is critical to preserving access to reliable information on the internet and preventing the spread of disinformation. Prior machine learning approaches rely on stylistic and meta-data features, but do not prioritise adaptability to author-specific behaviours. As a result, they struggle to effectively model the behaviour of specific sockpuppet-groups, especially when text data is limited. To address this, we propose the application of meta-learning, a machine learning technique designed to improve performance in data-scarce settings by training models across multiple tasks. Meta-learning optimises a model for rapid adaptation to the writing style of a new sockpuppet-group. Our results show that meta-learning significantly enhances the precision of predictions compared to pre-trained models, marking an advancement in combating sockpuppetry on open editing platforms. We release a new dataset of sockpuppet investigations to foster future research in both sockpuppetry and meta-learning fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。