对比多种模型,发现简单模型更适合作为临床决策支持工具。
Scalable Clinical Data Infrastructure and Comparative ML Evaluation for Hospitalisation Risk Prediction in Elderly Patients with Multiple Long-Term Conditions using CPRD

- 构建了基于CPRD Aurum的老年人长期病患者数据流水线,涵盖260种疾病。
- 在12个月急诊住院风险预测中,LASSO模型校准性能最佳(斜率0.817)。
- 强调模型可解释性与校准度对临床部署的重要性,非仅追求高判别力。
深度学习模型在电子健康记录患者轨迹建模中日益流行,但其在真实临床环境中的优势很少经过严谨实证检验。本文针对CPRD Aurum中的老年多病患者,构建了完整的患者时间线处理流程,采用三级自动化框架识别260种临床状况,其中对17种复杂疾病设有专用检测逻辑。在此基础上,比较了时序图卷积神经网络(TG-CNN)、带LASSO正则化的逻辑回归和随机森林在预测12个月内全因急诊住院风险上的表现。交叉验证下,TG-CNN平均AUC-ROC略高于LASSO(0.712 vs. 0.705),但在保留测试集上,LASSO表现最优(AUC-ROC 0.733),优于随机森林(0.710)和TG-CNN(0.702)。然而,仅LASSO经Platt校准后达到可接受的校准斜率(0.817),随机森林(0.759)与TG-CNN(0.391)仍严重校准偏差。因此,尽管非最高判别力,但LASSO是更适合直接临床部署的模型。研究提出机器学习与医疗领域在数据基础设施、模型选择及校准与可解释性价值方面的关键启示。
原文摘要 · Abstract (English)
Deep learning architectures are increasingly proposed for patient trajectory modeling in electronic health records (EHRs), yet their advantage over simpler, more interpretable models is rarely subjected to rigorous empirical scrutiny in real-world clinical settings. We present a comprehensive patient timeline pipeline applied to elderly patients in CPRD Aurum, incorporating 260 clinical conditions classified via a three-tier automated framework including specialised detection logic for 17 complex conditions. Using this infrastructure, we benchmark Temporal Graph Convolutional Neural Networks (TG-CNN) against Logistic Regression with LASSO regularisation and Random Forests for predicting 12-month all-cause emergency hospitalisation risk, motivated by (but not filtered to) the elevated risk of adverse drug reactions. Under cross-validation, TG-CNN achieves a marginally higher mean AUC-ROC than LASSO (0.712 vs. 0.705), whereas on the held-out test set LASSO achieves the highest discrimination of three models (AUC-ROC 0.733, versus 0.710 for Random Forest and 0.702 for TG-CNN). We show, that discrimination alone is an incomplete criterion for clinical deployment: after Platt calibration, LASSO is the only model with an acceptable calibration slope (0.817), while Random Forest (0.759) and, TG-CNN (0.391) remain substantially miscalibrated. We argue that LASSO, not the highest-discriminating model, is the model best suited to direct clinical deployment. We present lessons for the machine learning and healthcare community regarding data infrastructure, model selection, and value of calibration and interpretability in high-stakes decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。