五种模型在肾病预测中内部表现完美,但外部部署时准确率暴跌。
Calibration, Uncertainty Communication, and Deployment Readiness in CKD Risk Prediction: A Framework Evaluation Study

- 用校准与置信区间方法评估模型可靠性
- 外部数据中预测准确率下降至0.48-0.58,覆盖率达不到目标
- 临床部署前必须检验模型在真实场景下的稳定性
针对慢性肾病(CKD)风险预测的机器学习模型常在内部测试集上获得完美的区分能力,但校准与不确定性量化却常被忽视,导致临床无法判断概率输出是否可信。本研究在UCI CKD数据集(400名患者,CKD患病率62.5%)上训练了五种分类器:逻辑回归、随机森林、XGBoost、带Platt缩放的SVM和高斯朴素贝叶斯。评估涵盖校准质量、分位数预测覆盖率及八项部署准备度指标。通过分布应力测试,将各模型最优校准版本应用于公开的MIMIC-IV demo队列(97名患者,CKD占比23.7%),以评估在患病率变化和特征缺失下的表现。使用期望校准误差(ECE)和布里尔得分衡量校准效果,采用分拆式合取预测实现90%边际覆盖率。所有模型在UCI测试集上达到AUROC 1.00;经等距重校准后内部ECE降至0.000–0.022。但在MIMIC-IV中,AUROC降至0.48–0.58,ECE升至0.68–0.76,覆盖率从0.80–0.98骤降至0.21–0.25(目标90%)。无一模型在部署准备度清单中得分超过4/16。内部表现优异并未迁移至外部环境,校准稳定性与合取覆盖性应在真实部署前验证。
原文摘要 · Abstract (English)
Machine learning models for chronic kidney disease (CKD) risk prediction often post strong discrimination scores on internal test sets. Calibration and uncertainty quantification get far less attention, leaving clinicians without reliable information about whether the probability outputs are accurate. We trained five classifiers on the UCI CKD dataset (400 patients, 62.5% CKD prevalence): logistic regression, random forest, XGBoost, SVM with Platt scaling, and Gaussian naive Bayes. We evaluated each across calibration quality, conformal prediction coverage, and an eight-criterion deployment readiness framework. A distributional stress-test applied the best-calibrated variant of each model to the open-access MIMIC-IV demo cohort (97 patients, 23.7% CKD) to assess behaviour under prevalence shift and feature missingness. We measured calibration before and after Platt scaling and isotonic regression using Expected Calibration Error and Brier Score, and quantified uncertainty through split conformal prediction targeting 90% marginal coverage. All five models reached AUROC 1.00 on the UCI test set. Isotonic recalibration reduced internal ECE to 0.000-0.022. On MIMIC-IV, AUROC fell to 0.48-0.58, ECE rose to 0.68-0.76, and conformal coverage dropped from 0.80-0.98 to 0.21-0.25 against a 90% target. No model scored above 4 out of 16 on the deployment readiness checklist. Near-perfect internal performance did not transfer. Calibration stability and conformal coverage should be evaluated on external data before any clinical prediction model moves toward deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。