综合195项研究,发现情感分析模型平均准确率80%,但报告方式常误导结果。
A meta-analysis on the performance of machine-learning based language models for sentiment analysis
- 基于12项研究特征,用三层随机效应模型分析模型表现
- 最优模型平均准确率达80%(95%CI: 0.76–0.84)
- 强调应标准化报告,如提供测试集混淆矩阵
本文开展了一项元分析,评估机器学习模型在推特数据情感分析中的表现。研究旨在估算平均性能,评估研究间与研究内异质性,并分析研究特征对模型性能的影响。依据PRISMA指南,从学术数据库中筛选出20项研究中的195项试验,涵盖12项研究特征。总体准确率是最常报告的指标,采用双反正弦变换和三层随机效应模型进行分析。经AIC优化的模型平均总体准确率为0.80(95%CI:0.76–0.84)。研究揭示两点关键发现:1)总体准确率虽普遍使用,但易受类别不平衡和情感类别数量影响,存在误导性,亟需标准化处理;2)模型性能应通过标准化报告(如独立测试集的混淆矩阵)进行可靠比较,而当前实践远未普及。
原文摘要 · Abstract (English)
This paper presents a meta-analysis evaluating ML performance in sentiment analysis for Twitter data. The study aims to estimate the average performance, assess heterogeneity between and within studies, and analyze how study characteristics influence model performance. Using PRISMA guidelines, we searched academic databases and selected 195 trials from 20 studies with 12 study features. Overall accuracy, the most reported performance metric, was analyzed using double arcsine transformation and a three-level random effects model. The average overall accuracy of the AIC-optimized model was 0.80 [0.76, 0.84]. This paper provides two key insights: 1) Overall accuracy is widely used but often misleading due to its sensitivity to class imbalance and the number of sentiment classes, highlighting the need for normalization. 2) Standardized reporting of model performance, including reporting confusion matrices for independent test sets, is essential for reliable comparisons of ML classifiers across studies, which seems far from common practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。