随机划分数据会影响模型精度,用区间估计实现公平评估
Variation in prediction accuracy due to randomness in data division and fair evaluation using interval estimation
- 用autoML构建33600个糖尿病模型,测试初始状态对结果的影响
- 发现模型精度呈依赖初始状态的分布,可近似正态分布
- 提出用区间估计法公平比较模型性能,适合临床研究者参考
本文针对使用机器学习构建预测模型时的一个基础问题展开研究。尽管已有大量基于大型队列研究和机器学习算法的疾病诊断与预测模型,但其泛化能力仍面临挑战。其中,数据集随机划分被认为是重要原因之一。本研究利用autoML框架和公开糖尿病数据,构建了33,600个依赖初始状态随机性的糖尿病诊断模型,并评估其预测精度。结果显示,预测精度呈现出依赖初始状态的分布特征。由于该分布近似服从正态分布,我们采用统计区间估计方法,估算预测精度的期望区间,以实现对预测模型精度的公平比较。
原文摘要 · Abstract (English)
This paper attempts to answer a "simple question" in building predictive models using machine learning algorithms. Although diagnostic and predictive models for various diseases have been proposed using data from large cohort studies and machine learning algorithms, challenges remain in their generalizability. Several causes for this challenge have been pointed out, and partitioning of the dataset with randomness is considered to be one of them. In this study, we constructed 33,600 diabetes diagnosis models with "initial state" dependent randomness using autoML (automatic machine learning framework) and open diabetes data, and evaluated their prediction accuracy. The results showed that the prediction accuracy had an initial state-dependent distribution. Since this distribution could follow a normal distribution, we estimated the expected interval of prediction accuracy using statistical interval estimation in order to fairly compare the accuracy of the prediction models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。