一个统一模型搞定MRI全流程,从重建到报告
OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- 用多阶段训练融合图像与文本信息,实现跨任务通用
- 在22万例影像上训练,支持重建、分割、诊断等8类任务
- 适合临床研发者和医学AI工程师快速构建端到端系统
磁共振成像(MRI)在临床中不可或缺,但其工作流程分散,涵盖采集、重建、分割、检测、诊断和报告等多个阶段。尽管深度学习在单个任务上取得进展,现有方法多局限于特定解剖结构或应用场景,缺乏跨场景泛化能力,且极少整合放射科医生依赖的文本信息。本文提出OmniMRI,一个统一的视觉-语言基础模型,旨在覆盖完整MRI工作流。模型基于60个公开数据集构建的大规模异构语料库训练,包含超过22万例MRI体积和1900万张切片,涵盖图像仅数据、图像-文本配对数据及指令-响应数据。采用多阶段训练策略:自监督视觉预训练、视觉-语言对齐、多模态预训练及多任务指令微调,逐步赋予模型可迁移的视觉表征、跨模态推理和强指令遵循能力。定性结果表明,OmniMRI可在单一架构中完成包括重建、解剖与病灶分割、异常检测、诊断建议和放射科报告生成在内的多种任务。该成果展示了将碎片化流程整合为可扩展通用框架的潜力,推动实现影像与临床语言统一的基础模型,实现全面、端到端的MRI解读。
原文摘要 · Abstract (English)
Magnetic Resonance Imaging (MRI) is indispensable in clinical practice but remains constrained by fragmented, multi-stage workflows encompassing acquisition, reconstruction, segmentation, detection, diagnosis, and reporting. While deep learning has achieved progress in individual tasks, existing approaches are often anatomy- or application-specific and lack generalizability across diverse clinical settings. Moreover, current pipelines rarely integrate imaging data with complementary language information that radiologists rely on in routine practice. Here, we introduce OmniMRI, a unified vision-language foundation model designed to generalize across the entire MRI workflow. OmniMRI is trained on a large-scale, heterogeneous corpus curated from 60 public datasets, over 220,000 MRI volumes and 19 million MRI slices, incorporating image-only data, paired vision-text data, and instruction-response data. Its multi-stage training paradigm, comprising self-supervised vision pretraining, vision-language alignment, multimodal pretraining, and multi-task instruction tuning, progressively equips the model with transferable visual representations, cross-modal reasoning, and robust instruction-following capabilities. Qualitative results demonstrate OmniMRI's ability to perform diverse tasks within a single architecture, including MRI reconstruction, anatomical and pathological segmentation, abnormality detection, diagnostic suggestion, and radiology report generation. These findings highlight OmniMRI's potential to consolidate fragmented pipelines into a scalable, generalist framework, paving the way toward foundation models that unify imaging and clinical language for comprehensive, end-to-end MRI interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。