arXiv:2410.20293cs.LGstat.ML2024-10综述被引 27

系统评估36项ML研究,揭示社交平台骗术检测的算法偏见与改进方向

A Systematic Review of Machine Learning Approaches for Detecting Deceptive Activities on Social Media: Methods, Challenges, and Biases

  • 梳理36篇论文,分析从数据到模型全周期的常见偏差
  • 发现准确率误用、语言预处理不足等问题在多数研究中存在
  • 建议改用F1、AUROC等指标,提升模型可靠性与可推广性

Twitter、Facebook和Instagram等社交平台助长了虚假信息传播,亟需自动化检测系统。本系统综述评估了36项应用机器学习(ML)与深度学习(DL)模型检测假新闻、垃圾信息及虚假账号的研究。基于预测模型偏倚评估工具(PROBAST),识别出全生命周期的关键偏差:采样不具代表性导致的选择偏差、类别不平衡处理不足、语言预处理不充分(如否定词处理缺失)、超参数调优不一致。尽管支持向量机(SVM)、随机森林和长短期记忆网络(LSTM)展现潜力,但多数研究在类别不平衡场景下过度依赖准确率作为评估指标。综述强调需改进数据预处理(如重采样技术)、统一超参数调优,并采用精确率、召回率、F1分数和AUROC等更合适指标。解决这些局限可提升模型可靠性与泛化能力,助力减少社交平台上的虚假信息。

原文摘要 · Abstract (English)

Social media platforms like Twitter, Facebook, and Instagram have facilitated the spread of misinformation, necessitating automated detection systems. This systematic review evaluates 36 studies that apply machine learning (ML) and deep learning (DL) models to detect fake news, spam, and fake accounts on social media. Using the Prediction model Risk Of Bias ASsessment Tool (PROBAST), the review identified key biases across the ML lifecycle: selection bias due to non-representative sampling, inadequate handling of class imbalance, insufficient linguistic preprocessing (e.g., negations), and inconsistent hyperparameter tuning. Although models such as Support Vector Machines (SVM), Random Forests, and Long Short-Term Memory (LSTM) networks showed strong potential, over-reliance on accuracy as an evaluation metric in imbalanced data settings was a common flaw. The review highlights the need for improved data preprocessing (e.g., resampling techniques), consistent hyperparameter tuning, and the use of appropriate metrics like precision, recall, F1 score, and AUROC. Addressing these limitations can lead to more reliable and generalizable ML/DL models for detecting deceptive content, ultimately contributing to the reduction of misinformation on social media.

机器学习虚假信息偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。