用真实与合成数据提升心脏病介入术后三年死亡预测精度
Cardiac mortality prediction in patients undergoing PCI based on real and synthetic data
- 通过生成500个合成样本缓解数据不平衡问题
- 模型在少数类召回率上显著提升,且保持高区分能力
- 年龄、射血分数等四项指标对死亡风险影响最大
患者状态、造影及操作特征包含术后长期预后的关键信号。本研究基于2,044例接受冠脉分叉病变支架术(PCI)患者的实测与合成数据,构建心脏死亡风险预测模型,主要终点为3年随访期间的心脏死亡。采用多种机器学习模型进行预测,并通过额外生成500个合成样本以解决类别不平衡问题。为评估特征贡献,应用排列特征重要性分析;另开展实验验证移除非信息特征后模型表现变化。未使用过采样时,各模型总体准确率均达0.92–0.93,但几乎忽略少数类。经数据增强后,所有模型在少数类召回率上持续提升,对数似然与校准性能改善,且在高危人群构造中给出更合理的风险估计。特征重要性分析显示,年龄、射血分数、外周动脉疾病和脑血管疾病为最核心影响因素。
原文摘要 · Abstract (English)
Patient status, angiographic and procedural characteristics encode crucial signals for predicting long-term outcomes after percutaneous coronary intervention (PCI). The aim of the study was to develop a predictive model for assessing the risk of cardiac death based on the real and synthetic data of patients undergoing PCI and to identify the factors that have the greatest impact on mortality. We analyzed 2,044 patients, who underwent a PCI for bifurcation lesions. The primary outcome was cardiac death at 3-year follow-up. Several machine learning models were applied to predict three-year mortality after PCI. To address class imbalance and improve the representation of the minority class, an additional 500 synthetic samples were generated and added to the training set. To evaluate the contribution of individual features to model performance, we applied permutation feature importance. An additional experiment was conducted to evaluate how the model's predictions would change after removing non-informative features from the training and test datasets. Without oversampling, all models achieve high overall accuracy (0.92-0.93), yet they almost completely ignore the minority class. Across models, augmentation consistently increases minority-class recall with minimal loss of AUROC, improves probability quality, and yields more clinically reasonable risk estimates on the constructed severe profiles. According to feature importance analysis, four features emerged as the most influential: Age, Ejection Fraction, Peripheral Artery Disease, and Cerebrovascular Disease. These results show that straightforward augmentation with realistic and extreme cases can expose, quantify, and reduce brittleness in imbalanced clinical prediction using only tabular records, and motivate routine reporting of probability quality and stress tests alongside headline metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。