用可解释正则化提升多发性骨髓瘤生存预测模型的可信度
Interpretable Multiple Myeloma Prognosis with Observational Medical Outcomes Partnership Data
- 设计两种正则化方法,强制模型依赖可解释特征或符合临床分期标准
- 在812例患者数据上测试,模型准确率达0.721,且重要特征与临床一致
- 适合关注医疗模型可解释性的研究者和临床决策支持系统开发者
机器学习有望改善临床决策,但模型不透明限制其在医疗领域的应用。本文提出两种新型正则化技术,确保基于真实世界数据训练的机器学习模型具备可解释性。以赫尔辛基大学医院的临床数据为基础,预测多发性骨髓瘤患者的五年生存率。为保证模型可解释性,采用两种不同的正则化惩罚项构造方式:第一种惩罚模型预测偏离人工选取两个特征的可解释逻辑回归结果;第二种要求模型预测与修订版国际分期系统(R-ISS)保持一致。通过812名患者的实验验证,所提方法在测试集上达到最高0.721的准确率,且SHAP值显示模型依赖于选定的重要特征。
原文摘要 · Abstract (English)
Machine learning (ML) promises better clinical decision-making, yet opaque model behavior limits the adoption in healthcare. We propose two novel regularization techniques for ensuring the interpretability of ML models trained on real-world data. In particular, we consider the prediction of five-year survival for multiple myeloma patients using clinical data from Helsinki University Hospital. To ensure the interpretability of the trained models, we use two alternative constructions for a penalty term used for regularization. The first one penalizes deviations from the predictions obtained from an interpretable logistic regression method with two manually chosen features. The second construction requires consistency of model predictions with the revised international staging system (R-ISS). We verify the usefulness of the proposed regularization techniques in numerical experiments using data from 812 patients. They achieve an accuracy up to 0.721 on a test set and SHAP values show that the models rely on the selected important features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。