用心电图和病历数据多模态预测心功能分级,辅助基层医疗筛查。
A Multimodal and Explainable Machine Learning Approach to Diagnosing Multi-Class Ejection Fraction from Electrocardiograms
- 融合心电图与电子病历特征,构建多模态分类模型。
- 在4类心功能分级中最高达到0.95的AUROC,优于单一数据源。
- 通过可解释性分析识别关键特征,适合临床辅助决策场景。
左心室射血分数(LVEF)评估依赖超声心动图,在基层医疗和资源有限地区难以普及。本文提出一种多模态机器学习框架,结合12导联心电图时序特征与结构化电子健康记录(EHR)变量,将LVEF分为四类:正常(>50%)、轻度降低(40-50%)、中度降低(30-40%)、重度降低(<30%)。为增强模型可解释性,采用SHAP方法识别最具影响力的ECG与EHR特征。基于哈特福德健康医疗系统回顾性数据,使用36,784对心电图-超声心动图数据训练XGBoost模型,并在后续19,966例心电图上验证其时间泛化能力。多模态模型在一对多分类中分别获得0.95(重度)、0.92(中度)、0.82(轻度)和0.91(正常)的AUROC,优于仅使用心电图或仅使用病历的基线模型,且在时间验证中表现稳定。该研究支持以心电图为依据的多模态分层策略,作为资源受限环境中优先确认影像检查的实用筛查工具。
原文摘要 · Abstract (English)
Left ventricular ejection fraction (LVEF) assessment depends on echocardiography, limiting access in primary care and resource-constrained settings. We developed a multimodal machine-learning framework that combines engineered 12-lead ECG timeseries features with structured EHR variables to classify LVEF into four clinically used strata: normal (>50%), mildly reduced (40-50%), moderately reduced (30-40%), and severely reduced (<30%). To support model explainability, we identified the most influential ECG and EHR features via SHAP attributions. Using retrospective data from Hartford HealthCare, we trained XGBoost models on 36,784 ECG-echocardiogram pairs from 30,952 outpatients and evaluated temporal generalizability on 19,966 ECGs from a subsequent period. The multimodal model achieved one-vs-rest AUROCs of 0.95 (severe), 0.92 (moderate), 0.82 (mild), and 0.91 (normal), outperforming ECG-only and EHR-only baselines, and maintained performance under temporal validation. This work supports ECG-based, multimodal LVEF stratification as a practical screening and triage aid to prioritize confirmatory imaging where resources are limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。