提升心脏MRI影像视图分类准确率,应对临床扫描差异挑战
ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines
- 结合3D卷积与多尺度自注意力,捕捉局部结构与心脏周期动态
- 在15万例数据上达96%准确率,各类F1超0.94,校准误差仅0.025
- 适合需高鲁棒性视图识别的临床端到端心脏影像分析系统
可靠识别标准心肌磁共振电影视图至关重要,因每个视图决定可视解剖结构及可进行的定量分析。人为或自动系统误判视图会将错误传递至分割、容积评估、应变分析和瓣膜评价。然而,在扫描设备、采集协议、运动伪影和成像平面设定等临床常规变异下,准确分类仍具挑战。本文提出ConvFormer3D-TAP,一种专为电影序列设计的时空架构,融合3D卷积标记化与多尺度自注意力。模型通过掩码时空重建和不确定性加权多片段融合训练,增强对心脏各阶段及模糊时序段的鲁棒性。该设计同时捕捉局部解剖结构(卷积先验)与长程心动周期动态(分层注意力)。在涵盖六种标准心肌磁共振电影视图的150,974例临床序列上,模型验证准确率达96%,各类F1-score ≥0.94,校准性能优异(ECE=0.025;Brier=0.040)。误差分析显示残余混淆集中于解剖邻近的长轴与左室流出道/主动脉瓣视图对,符合固有平面重叠特性。结果表明ConvFormer3D-TAP可作为端到端心脏磁共振工作流中视图路由、过滤与质量控制的可扩展前端。
原文摘要 · Abstract (English)
Reliable recognition of standard cine cardiac MRI views is essential because each view determines which cardiac anatomy is visualized and which quantitative analyses can be performed. Incorrect view identification, whether by a human reader or an automated deep learning system, can propagate errors into segmentation, volumetric assessment, strain analysis, and valve evaluation. However, accurate view classification remains challenging under routine clinical variability in scanner vendor, acquisition protocol, motion artifacts, and plane prescription. We present ConvFormer3D-TAP, a cine-specific spatiotemporal architecture that integrates 3D convolutional tokenization with multiscale self-attention. The model is trained using masked spatiotemporal reconstruction and uncertainty-weighted multi-clip fusion to enhance robustness across cardiac phases and ambiguous temporal segments. The design captures complementary cues: local anatomical structure through convolutional priors and long-range cardiac-cycle dynamics through hierarchical attention. On a cohort of 150,974 clinically acquired cine sequences spanning six standard cine cardiac MRI views, ConvFormer3D-TAP achieved 96% validation accuracy with per-class F1-scores >= 0.94 and strong calibration (ECE = 0.025; Brier = 0.040). Error analysis shows that residual confusions are concentrated in anatomically adjacent long-axis and LVOT/AV view pairs, consistent with intrinsic prescription overlap. These results support ConvFormer3D-TAP as a scalable front-end for view routing, filtering and quality control in end-to-end cMRI workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。