用深度学习自动分割心脏磁共振图像,提升疾病检测精度。
Deep learning-based segmentation of T1 and T2 cardiac MRI maps for automated disease detection
- 用深度学习模型自动分割心肌和血池,准确率超人工标注差异。
- 结合均值、四分位数等多特征,分类F1分数达92.7%。
- 适合心血管影像分析、医学人工智能研究者参考。
参数化组织映射可实现心肌定量分析,但手动分割存在观察者间差异。传统方法依赖平均弛豫值与单一阈值,难以捕捉心肌复杂性。本研究评估深度学习(DL)是否能达到与观察者间一致性相当的分割精度,探索除平均T1/T2值外的统计特征价值,并检验融合多特征的机器学习(ML)在疾病检测中的效果。对T1和T2图进行人工分割,测试集由两名观察者独立标注以评估观察者间差异。训练深度学习模型分割左心室血池与心肌。计算心肌像素的平均值(A)、下四分位数(LQ)、中位数(M)和上四分位数(UQ),用于设定阈值或输入机器学习分类器。采用骰子相似系数(DICE)和平均绝对百分比误差评估分割性能,Bland-Altman图分析观察者与模型间一致性,受试者工作特征分析确定最优阈值,皮尔逊相关性比较模型与人工分割特征,F1-score、精确率和召回率评估分类性能。采用威尔科克斯检验比较方法差异,p < 0.05认为有统计学意义。共纳入144名受试者,分为训练(100人)、验证(15人)和评估(29人)组。分割模型取得85.4% DICE,优于观察者间一致性。随机森林结合所有特征使F1-score提升至92.7%(p < 0.001)。结论:深度学习可实现T1/T2图高效分割,融合多特征的机器学习显著提升疾病检测能力。
原文摘要 · Abstract (English)
Objectives Parametric tissue mapping enables quantitative cardiac tissue characterization but is limited by inter-observer variability during manual delineation. Traditional approaches relying on average relaxation values and single cutoffs may oversimplify myocardial complexity. This study evaluates whether deep learning (DL) can achieve segmentation accuracy comparable to inter-observer variability, explores the utility of statistical features beyond mean T1/T2 values, and assesses whether machine learning (ML) combining multiple features enhances disease detection. Materials & Methods T1 and T2 maps were manually segmented. The test subset was independently annotated by two observers, and inter-observer variability was assessed. A DL model was trained to segment left ventricle blood pool and myocardium. Average (A), lower quartile (LQ), median (M), and upper quartile (UQ) were computed for the myocardial pixels and employed in classification by applying cutoffs or in ML. Dice similarity coefficient (DICE) and mean absolute percentage error evaluated segmentation performance. Bland-Altman plots assessed inter-user and model-observer agreement. Receiver operating characteristic analysis determined optimal cutoffs. Pearson correlation compared features from model and manual segmentations. F1-score, precision, and recall evaluated classification performance. Wilcoxon test assessed differences between classification methods, with p < 0.05 considered statistically significant. Results 144 subjects were split into training (100), validation (15) and evaluation (29) subsets. Segmentation model achieved a DICE of 85.4%, surpassing inter-observer agreement. Random forest applied to all features increased F1-score (92.7%, p < 0.001). Conclusion DL facilitates segmentation of T1/ T2 maps. Combining multiple features with ML improves disease detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。