为官方统计中的机器学习应用建立误差分析框架,提升结果可信度。
Leveraging Machine Learning for Official Statistics: A Statistical Manifesto
- 提出总机器学习误差(TMLE)框架,类比调查误差模型。
- 系统识别并量化模型内部与外部有效性问题。
- 适合政策制定者、统计机构和机器学习从业者参考。
官方统计生产应用机器学习需具备统计严谨性,尽管近年来技术飞速发展,但其方法论稳健性仍不足以产出高质量统计结果。为全面考量机器学习模型的所有误差来源,本文提出总机器学习误差(TMLE)框架,类比于调查方法中的总调查误差模型。该框架旨在确保机器学习模型在内部与外部均具有效性,涵盖代表性偏差与测量误差等问题。文中通过多个案例研究,说明在官方统计中更严格地应用机器学习的必要性。
原文摘要 · Abstract (English)
It is important for official statistics production to apply ML with statistical rigor, as it presents both opportunities and challenges. Although machine learning has enjoyed rapid technological advances in recent years, its application does not possess the methodological robustness necessary to produce high quality statistical results. In order to account for all sources of error in machine learning models, the Total Machine Learning Error (TMLE) is presented as a framework analogous to the Total Survey Error Model used in survey methodology. As a means of ensuring that ML models are both internally valid as well as externally valid, the TMLE model addresses issues such as representativeness and measurement errors. There are several case studies presented, illustrating the importance of applying more rigor to the application of machine learning in official statistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。