临床风险模型中,常用数据平衡方法反而降低预测准确性。
Tipping the Balance: Impact of Class Imbalance Correction on the Performance of Clinical Risk Prediction Models
- 对比原始数据与三种数据重采样策略,评估模型表现
- 所有方法均未提升判别能力,且多数降低校准精度
- 尤其适合关注概率准确性的医疗决策研究者阅读
目标:基于机器学习的临床风险预测模型广泛用于医疗决策支持。尽管类别不平衡修正技术常用于稀有事件场景以提升性能,但其对概率校准的影响仍不明确。本研究评估了多种常见重采样策略在真实世界临床预测任务中对判别与校准的影响。方法:分析了涵盖10个不同医学领域的临床数据集,共605,842名患者。评估了线性与多种非线性机器学习模型家族。模型在原始数据及三种常见1:1不平衡修正策略(SMOTE、RUS、ROS)下训练,使用保留数据集上的判别与校准指标进行评估。结果:在所有数据集和模型族中,重采样对预测性能无正面影响。与原始数据训练模型相比,受试者工作特征曲线下面积(ROC-AUC)变化微小且不一致(ROS: -0.002, p<0.05;RUS: -0.004, p>0.05;SMOTE: -0.01, p<0.05),无任一策略系统性提升。相反,重采样普遍损害校准性能。使用不平衡修正训练的模型表现出更高的布里尔分数(0.029至0.080, p<0.05),反映概率准确性下降,并出现显著的校准截距与斜率偏差,表明预测风险系统性扭曲,尽管排序性能保持不变。结论:在多样化的现实临床预测任务中,常用的类别不平衡修正技术未能带来可推广的判别性能提升,且伴随校准性能下降。
原文摘要 · Abstract (English)
Objective: ML-based clinical risk prediction models are increasingly used to support decision-making in healthcare. While class-imbalance correction techniques are commonly applied to improve model performance in settings with rare outcomes, their impact on probabilistic calibration remains insufficiently understood. This study evaluated the effect of widely used resampling strategies on both discrimination and calibration across real-world clinical prediction tasks. Methods: Ten clinical datasets spanning diverse medical domains and including 605,842 patients were analyzed. Multiple machine-learning model families, including linear models and several non-linear approaches, were evaluated. Models were trained on the original data and under three commonly used 1:1 class-imbalance correction strategies (SMOTE, RUS, ROS). Performance was assessed on held-out data using discrimination and calibration metrics. Results: Across all datasets and model families, resampling had no positive impact on predictive performance. Changes in the Receiver Operating Characteristic Area Under Curve (ROC-AUC) relative to models trained on the original data were small and inconsistent (ROS: -0.002, p<0.05; RUS: -0.004, p>0.05; SMOTE: -0.01, p<0.05), with no resampling strategy demonstrating a systematic improvement. In contrast, resampling in general degraded the calibration performance. Models trained using imbalance correction exhibited higher Brier scores (0.029 to 0.080, p<0.05), reflecting poorer probabilistic accuracy, and marked deviations in calibration intercept and slope, indicating systematic distortions of predicted risk despite preserved rank-based performance. Conclusion: In a diverse set of real-world clinical prediction tasks, commonly used class-imbalance correction techniques did not provide generalizable improvements in discrimination and were associated with degraded calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。