构建3D MRI通用模型,提升医学影像智能分析能力
Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
- 用20万组MRI数据训练,融合视觉与报告文本监督
- 在疾病分类等任务上优于现有模型,提升显著
- 模块化设计适合临床研究快速适配新任务
磁共振成像在临床诊断与研究中至关重要,但其复杂性和异质性限制了机器学习的可扩展性与泛化能力。尽管基础模型已革新自然语言与视觉任务,其在MRI领域的应用仍受限于数据稀缺与解剖范围狭窄。我们提出Decipher-MR,一种面向3D MRI的视觉-语言基础模型,基于超过22,000例研究中的200,000个MRI序列训练,覆盖多样解剖区域、成像序列和病灶类型。Decipher-MR结合自监督视觉学习与报告引导的文本监督,构建适用于广泛场景的鲁棒表征。为支持高效应用,该模型采用模块化设计,允许在冻结预训练编码器基础上微调轻量级任务专用解码器。在此设置下,我们在疾病分类、人口统计预测、解剖定位及跨模态检索任务中评估,结果表明其性能持续优于现有基础模型与专用方法。这些成果使Decipher-MR成为临床与科研中基于MRI的人工智能通用基础。
原文摘要 · Abstract (English)
Magnetic Resonance Imaging is a critical imaging modality in clinical diagnosis and research, yet its complexity and heterogeneity hinder scalable, generalizable machine learning. Although foundation models have revolutionized language and vision tasks, their application to MRI remains constrained by data scarcity and narrow anatomical focus. We present Decipher-MR, a 3D MRI-specific vision-language foundation model trained on 200,000 MRI series from over 22,000 studies spanning diverse anatomical regions, sequences, and pathologies. Decipher-MR integrates self-supervised vision learning with report-guided text supervision to build robust representations for broad applications. To enable efficient use, Decipher-MR supports a modular design that enables tuning of lightweight, task-specific decoders attached to a frozen pretrained encoder. Following this setting, we evaluate Decipher-MR across disease classification, demographic prediction, anatomical localization, and cross-modal retrieval, demonstrating consistent improvements over existing foundation models and task-specific approaches. These results position Decipher-MR as a versatile foundation for MRI-based AI in clinical and research settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。