评估遥感机器学习模型对可持续发展目标的监测性能,发现准确率易受数据不平衡影响。
Performance of models for monitoring sustainable development goals from remote sensing: A three-level meta-regression
- 采用三层随机效应模型分析86项研究的性能表现
- 最佳模型平均准确率达90%(95%置信区间0.86~0.92)
- 建议标准化报告混淆矩阵以提升跨研究可比性
机器学习(ML)是利用遥感数据监测联合国可持续发展目标(SDGs)的重要工具。本文通过元分析评估了ML在遥感数据中用于监测SDGs的性能表现,旨在:1)估算平均性能;2)确定研究间与研究内异质性程度;3)评估研究特征对模型性能的影响。基于PRISMA指南,在多个学术数据库中检索相关研究,经三名评审员筛选200篇文献后,最终纳入20项研究中的86个试验,涵盖14项研究特征。总体准确率是最常报告的性能指标,采用双反正弦变换和三层随机效应模型进行分析。最佳模型的平均总体准确率为0.90(95%置信区间[0.86, 0.92])。模型性能存在显著异质性,其中64%来源于研究间差异。唯一显著影响因素是多数类样本比例,解释了61%的研究间异质性,其余13项特征均未显著提升模型表现。主要贡献在于两点:1)总体准确率虽最常用,但对类别不平衡敏感,需归一化处理,而该做法远未普及;2)领域亟需统一报告规范,独立测试集的混淆矩阵报告是实现模型分类器跨研究比较的关键。这些发现强调了在机器学习应用中建立稳健、可比评估指标的重要性,以确保可靠的可持续发展目标监测与政策制定。
原文摘要 · Abstract (English)
Machine learning (ML) is a tool to exploit remote sensing data for the monitoring and implementation of the United Nations' Sustainable Development Goals (SDGs). In this paper, we report on a meta-analysis to evaluate the performance of ML applied to remote sensing data to monitor SDGs. Specifically, we aim to 1) estimate the average performance; 2) determine the degree of heterogeneity between and within studies; and 3) assess how study features influence model performance. Using PRISMA guidelines, a search was performed across multiple academic databases to identify potentially relevant studies. A random sample of 200 was screened by three reviewers, resulting in 86 trials within 20 studies with 14 study features. Overall accuracy was the most reported performance metric. It was analyzed using double arcsine transformation and a three-level random effects model. The average overall accuracy of the best model was 0.90 [0.86, 0.92]. There was considerable heterogeneity in model performance, 64% of which was between studies. The only significant feature was the prevalence of the majority class, which explained 61% of the between-study heterogeneity. None of the other thirteen features added value to the model. The most important contributions of this paper are the following two insights. 1) Overall accuracy is the most popular performance metric, yet arguably the least insightful. Its sensitivity to class imbalance makes it necessary to normalize it, which is far from common practice. 2) The field needs to standardize the reporting. Reporting of the confusion matrix for independent test sets is the most important ingredient for between-study comparisons of ML classifiers. These findings underscore the need for robust and comparable evaluation metrics in machine learning applications to ensure reliable and actionable insights for effective SDG monitoring and policy formulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。