综合81种机器学习方法,发现信息欺诈检测平均准确率超79%。
Truth in Text: A Meta-Analysis of ML-Based Cyber Information Influence Detection Approaches
- 通过两阶段元分析整合81项研究,评估机器学习检测虚假信息效果。
- 整体平均准确率达79.18%,多数方法超过80%。
- 不同模型类型效果无显著差异,但组内差异大,适合复现与优化研究。
网络信息影响(即虚假信息)被视为威胁社会进步与政府稳定的主要因素。从美国总统选举到欧盟公投,乃至区域性山火新闻报道,虚假信息已常态化并扭曲决策。为此,学界涌现出大量旨在检测在线媒体中虚假信息的研究。前沿研究采用多种机器学习技术,包括支持向量机、随机森林、朴素贝叶斯等传统算法,以及卷积神经网络、长短期记忆网络和基于Transformer的深度模型。尽管总体成效显著,但文献整体存在不一致现象,限制了对真实有效性的理解。本研究采用两阶段元分析:(a) 计算机器学习模型在虚假信息检测中的总体效果统计;(b) 按模型类型分组分析。结果显示,81种检测方法中多数准确率高于80%,整体样本平均准确率为79.18%。各模型子组间无显著差异,但组内方差较大。研究建议未来工作应聚焦于模型层级的可复现性与方法改进。
原文摘要 · Abstract (English)
Cyber information influence, or disinformation in general terms, is widely regarded as one of the biggest threats to social progress and government stability. From US presidential elections to European Union referendums and down to regional news reporting of wildfires, lies and post-truths have normalized radical decision-making. Accordingly, there has been an explosion in research seeking to detect disinformation in online media. The frontier of disinformation detection research is leveraging a variety of ML techniques such as traditional ML algorithms like Support Vector Machines, Random Forest, and Naïve Bayes. Other research has applied deep learning models including Convolutional Neural Networks, Long Short-Term Memory networks, and transformer-based architectures. Despite the overall success of such techniques, the literature demonstrates inconsistencies when viewed holistically which limits our understanding of the true effectiveness. Accordingly, this work employed a two-stage meta-analysis to (a) demonstrate an overall meta statistic for ML model effectiveness in detecting disinformation and (b) investigate the same by subgroups of ML model types. The study found the majority of the 81 ML detection techniques sampled have greater than an 80\% accuracy with a Mean sample effectiveness of 79.18\% accuracy. Meanwhile, subgroups demonstrated no statistically significant difference between-approaches but revealed high within-group variance. Based on the results, this work recommends future work in replication and development of detection methods operating at the ML model level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。