一个可处理心脏影像多种任务的通用模型,性能超越传统方法。
A versatile foundation model for cine cardiac magnetic resonance image analysis tasks
- 用多视角混合架构在1500万张影像上训练,支持分割、定位、诊断等任务。
- 在4500+张独立影像上表现优于传统CNN,半量微调仍保持高精度。
- 不仅能判心脏病,还能预测寿命和发现系统性疾病影响,公平性好。
本文提出一种通用基础模型CineMA,可执行包括分割、关键点定位、诊断和预后在内的多种临床相关心脏影像分析任务。该模型基于多视图卷积-变压器掩码自编码器,在74,916名受试者的1500万张动态心脏磁共振影像上训练。在包含8个独立数据集、超过4500张影像的多任务验证中,其表现优于现有模型,是迄今为止最大规模的心脏磁共振影像基准研究。CineMA在心室边界勾画和射血分数估算(心脏功能关键指标)方面持续优于传统卷积神经网络(CNN),即使仅使用一半微调数据也保持优势。在疾病检测上超越CNN,长轴功能测量表现相当。有趣的是,该模型还可识别糖尿病、高血压、癌症等系统性疾病引起的心脏变化,并预测死亡风险。最后,模型公平性评估显示其在不同人口学亚组间表现一致。这些结果凸显CineMA的准确性、学习效率、适应性和公平性,具有支撑临床工作流与心血管研究的潜力。所有训练与推理代码及模型已公开于https://github.com/mathpluscode/CineMA。
原文摘要 · Abstract (English)
Here we present a versatile foundation model that can perform a range of clinically-relevant image analysis tasks, including segmentation, landmark localisation, diagnosis, and prognostication. A multi-view convolution-transformer masked autoencoder, named as CineMA, was trained on 15 million cine images from 74,916 subjects. The model was validated on multiple image analysis tasks and compared to existing models on >4,500 images from eight independent datasets with diverse population characteristics, representing the largest benchmark study for cine CMR so far. CineMA consistently outperformed conventional convolutional neural networks (CNNs) in delineating ventricular boundaries and estimating ejection fraction, a key measure of cardiac function. The improved performance was preserved, even when the model only used half of fine-tuning data. CineMA also surpassed CNNs in disease detection and matched their performance in long-axis function measurement. Interestingly, we found that CineMA can also detect cardiac changes in systemic diseases, such as diabetes, hypertension and cancer, and can also predict mortality. Finally, we assessed model fairness and demonstrated consistent model performance across demographic subgroups. These findings highlight CineMA's accuracy, learning efficiency, adaptability, and fairness, underscoring its potential as a foundation model for automated cardiac image analysis to support clinical workflow and cardiovascular research. All training and inference code and models are made publicly available at https://github.com/mathpluscode/CineMA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。