用机器学习补全数据后,仍能保持统计推断的高效与准确。
Another look at statistical inference with machine learning-imputed data
- 基于两阶段抽样思想,设计新方法处理机器学习预测数据
- 无论预测精度如何,效率都不低于传统纯金标准方法
- 适合生物医学等依赖低成本预测数据的科研场景
从结构生物学到流行病学,机器学习模型的预测正越来越多地补充昂贵的金标准数据,使科学研究更快、更便宜、更可扩展。为此,基于预测(PB)推断应运而生,旨在结合大量预测数据与少量金标准数据进行统计分析。其目标是双重的:(i) 减少由预测误差带来的偏差;(ii) 相较于仅使用金标准数据的经典推断,提升效率。尽管早期方法主要关注偏差缓解,效率提升仍是研究热点。受统计学及相关领域长期问题启发,我们借鉴两阶段抽样理论,提出一种针对机器学习插补结果的Z估计方法,该方法在任何预测质量下均保证效率不低于经典推断。通过理论和数值分析以及对英国生物样本库(UK Biobank)数据的应用验证了其有效性,并揭示了现有PB推断方法与经典及现代统计方法之间的新联系。
原文摘要 · Abstract (English)
From structural biology to epidemiology, predictions from machine learning (ML) models increasingly complement costly gold-standard data, enabling faster, more affordable, and scalable scientific inquiry. In response, prediction-based (PB) inference has emerged to support statistical analysis that combines a large volume of predicted data with a small amount of gold-standard data. The goals of PB inference are twofold: (i) to mitigate bias arising from prediction error and (ii) to improve efficiency relative to classical inference based solely on gold-standard data. While early PB inference methods primarily focused on bias mitigation, improving efficiency remains an active area of research. Motivated by connections between PB inference and longstanding problems in statistics and related fields, we draw on the two-phase sampling literature to introduce an approach for Z-estimation with ML-imputed outcomes that is guaranteed to match or exceed the efficiency of classical inference, regardless of prediction quality. We demonstrate the utility of our approach through theoretical and numerical analyses as well as an application to UK Biobank data. We further establish new connections between existing PB inference approaches and foundational and contemporary statistical methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。