轻量级模型实时精准估算超声心动图射血分数
Echo-E$^3$Net: Efficient Endocardial Spatio-Temporal Network for Ejection Fraction Estimation
- 结合心脏时相特异性边界检测与特征融合,建模心室壁运动
- 在EchoNet-Dynamic上实现5.20的RMSE、0.82的R²,仅需155万参数
- 无需预训练或复杂后处理,适合床旁超声快速部署
为开发一种鲁棒且计算高效的深度学习模型,用于从超声心动图视频中自动估计左心室射血分数(LVEF),以支持实时床旁超声(POCUS)部署。本文提出Echo-E$^3$Net,一种显式融入心脏解剖先验的内膜时空网络。模型包含双时相内膜边界检测器(E$^2$CBD),利用时相特异性交叉注意力定位收缩末期和舒张末期内膜标志点,并学习时相感知的标志点嵌入;以及内膜特征聚合器(E$^2$FA),将这些嵌入与深层特征图的全局统计描述符融合以优化EF回归。训练采用受辛普森双平面法启发的多组件损失,联合监督EF值与标志点几何结构。在EchoNet-Dynamic数据集上评估,使用RMSE与R$^2$指标,并报告参数量与GFLOPs以表征效率。结果表明,该模型在测试集上达到5.20的RMSE与0.82的R$^2$,仅消耗1.55M参数与8.05 GFLOPs,且无需外部预训练、大规模数据增强或测试时集成,支持实际实时部署。
原文摘要 · Abstract (English)
Objective To develop a robust and computationally efficient deep learning model for automated left ventricular ejection fraction (LVEF) estimation from echocardiography videos that is suitable for real-time point-of-care ultrasound (POCUS) deployment. Methods We propose Echo-E$^3$Net, an endocardial spatio-temporal network that explicitly incorporates cardiac anatomy into LVEF prediction. The model comprises a dual-phase Endocardial Border Detector (E$^2$CBD) that uses phase-specific cross attention to localize end-diastolic and end-systolic endocardial landmarks and to learn phase-aware landmark embeddings, and an Endocardial Feature Aggregator (E$^2$FA) that fuses these embeddings with global statistical descriptors of deep feature maps to refine EF regression. Training is guided by a multi-component loss inspired by Simpson's biplane method that jointly supervises EF and landmark geometry. We evaluate Echo-E$^3$Net on the EchoNet-Dynamic dataset using RMSE and R$^2$ while reporting parameter count and GFLOPs to characterize efficiency. Results On EchoNet-Dynamic, Echo-E$^3$Net achieves an RMSE of 5.20 and an R$^2$ score of 0.82 while using only 1.55M parameters and 8.05 GFLOPs. The model operates without external pre-training, heavy data augmentation, or test-time ensembling, supporting practical real-time deployment. Conclusion By combining phase-aware endocardial landmark modeling with lightweight spatio-temporal feature aggregation, Echo-E$^3$Net improves the efficiency and robustness of automated LVEF estimation and is well-suited for scalable clinical use in POCUS settings. Code is available at https://github.com/moeinheidari7829/Echo-E3Net
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。