用深度学习自动估测心脏超声视频的射血分数,准确率达93.2%
Investigating Deep Learning Models for Ejection Fraction Estimation from Echocardiography Videos
- 采用改进的3D Inception模型处理超声视频,融合时空特征
- 在EchoNet-Dynamic数据集上实现6.79%的均方根误差
- 揭示模型复杂度与泛化能力的权衡,适合医学影像自动化分析
左心室射血分数(LVEF)是评估心脏功能的关键指标,在心血管疾病诊断与管理中具有核心作用。超声心动图作为一种便捷且无创的影像手段,广泛用于临床LVEF估算。然而,人工评估超声心动图耗时且存在显著观察者间差异。深度学习提供了有前景的替代方案,有望达到经验丰富的医生水平。本研究系统评估了多种深度学习架构在超声视频中LVEF估计的表现,包括3D Inception、双流网络和CNN-RNN模型。通过调整网络结构与融合策略,优化预测精度。模型在包含10,030段超声视频的EchoNet-Dynamic数据集上训练与评估。结果表明,改进的3D Inception模型表现最佳,均方根误差(RMSE)为6.79%。各模型普遍存在过拟合现象,较小型简单模型通常具备更强泛化能力。模型性能对超参数极为敏感,尤其是卷积核大小与归一化策略。尽管聚焦于超声心动图,但关于架构设计与训练策略的发现可推广至更广泛的医学与非医学视频分析任务。
原文摘要 · Abstract (English)
Left ventricular ejection fraction (LVEF) is a key indicator of cardiac function and plays a central role in the diagnosis and management of cardiovascular disease. Echocardiography, as a readily accessible and non-invasive imaging modality, is widely used in clinical practice to estimate LVEF. However, manual assessment of cardiac function from echocardiograms is time-consuming and subject to considerable inter-observer variability. Deep learning approaches offer a promising alternative, with the potential to achieve performance comparable to that of experienced human experts. In this study, we investigate the effectiveness of several deep learning architectures for LVEF estimation from echocardiography videos, including 3D Inception, two-stream, and CNN-RNN models. We systematically evaluate architectural modifications and fusion strategies to identify configurations that maximize prediction accuracy. Models were trained and evaluated on the EchoNet-Dynamic dataset, comprising 10,030 echocardiogram videos. Our results demonstrate that modified 3D Inception architectures achieve the best overall performance, with a root mean squared error (RMSE) of 6.79%. Across architectures, we observe a tendency toward overfitting, with smaller and simpler models generally exhibiting improved generalization. Model performance was also found to be highly sensitive to hyperparameter choices, particularly convolutional kernel sizes and normalization strategies. While this study focuses on echocardiography-based LVEF estimation, the insights gained regarding architectural design and training strategies may be applicable to a broader range of medical and non-medical video analysis tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。